Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Physical Mechanism behind the Early Onset of the Ultimate State in Supergravitational Centrifugal Thermal Convection
arXiv:2509.14800v2 Announce Type: replace Abstract: We present a combined experimental and numerical investigation of the transition from the classical to the ultimate regime of thermal turbulence in a supergravitational centrifugal convection system. The transition is found to be robust, with the critical Rayleigh number decreasing systematically as the Froude number, defined as the ratio of centrifugal to Earth's gravity, decreases, highlighting the effect of residual gravity. Once the Rayleigh number reaches the transition threshold, the Stewartson layer induced by residual Earth gravity becomes comparable in thickness to the viscous boundary layer, and their interaction results in a coupled flow that distorts the viscous boundary layer, triggering its transition from laminar to turbulent flow and leading to a sharp increase in heat transport. These findings demonstrate the key role of the Stewartson layer induced by residual gravity in facilitating the transition to the ultimate regime in supergravitational centrifugal thermal convection.
AI Behavioral Science
arXiv:2509.13323v2 Announce Type: replace Abstract: We outline a foundation for a new field of ``AI Behavioral Science,'' covering three perspectives. First, as AI becomes ubiquitous and is increasingly proprietary and opaque, it becomes vital to develop techniques for assessing AI behavior. We outline how tools developed to assess people's behaviors by social scientists can be used to assess and infer AI's behaviors biases, tendencies, and heuristics. Second, we also discuss how AI can change the ways in which we learn about human behavior. Beyond its computational power, AI offers new techniques for simulating, inferring, and predicting human behaviors that we outline and discuss. Third, as humans and AI are interacting in increasingly complex and intertwined systems, we need to understand the implications for the resulting economic and political outcomes. We outline issues that are increasingly pressing concerning the future of human-AI interactions and potential changes and disruptions that can ensue.
Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning
arXiv:2605.31174v1 Announce Type: new Abstract: Object detection in real-world scenarios remains challenging due to diverse image degradations and heterogeneous object distributions, which significantly hinder the generalization of existing detectors. Conventional approaches, including scene-specific representation learning and end-to-end pipeline design, are inherently limited by their reliance on predefined conditions and lack adaptability to dynamic environments. In this paper, we propose DetAS, an agentic detection framework that formulates object detection as a dynamic decision process. Instead of relying on static pipelines, DetAS leverages a Multimodal Large Language Model (MLLM) as a central agent to adaptively compose detection workflows by selecting from a toolbox of restoration modules and specialized detectors. Specifically, DetAS consists of two key components: Self-Adaptive Image Restoration, which dynamically determines whether and how to enhance images for downstream detection, and Multi-Expertise Detection, which integrates multiple domain-specialized detectors and resolves their predictions through instance-level reasoning. To further improve decision quality under fine-grained conditions, we introduce Self-Evolving Experience Harvesting and extend the framework to DetAS-X, which accumulates node-level decision experience from a small set of annotated data and enables experience-aware reasoning during inference. This mechanism allows the system to progressively refine its decision policy and adapt to diverse real-world scenarios. Extensive experiments on six challenging benchmarks demonstrate that DetAS-X significantly outperforms existing MLLM-based detectors, achieving an average improvement of 28.36% in F1 score, with up to 37.01% gain on DarkFace. These results demonstrate the promise of agentic detection and establish a solid foundation for its application in complex and dynamic environments.
TUX: Measuring Human--AI Tacit Understanding
arXiv:2605.30930v1 Announce Type: new Abstract: As large language models (LLMs) increasingly act as collaborative partners, human--AI alignment is often evaluated through explicit task success, accuracy, or reward optimization. Yet many collaborative settings depend on tacit understanding: whether an agent can align with a human's evaluative stance or representational priors without clear objectives, communication, or feedback. To study this capacity, we develop a spectrum-placement task inspired by the social party game Wavelength, in which humans and agents independently place concepts along subjective spectra. We operationalize the Tacit Understanding Index (TUX) as a pairwise measure of similarity between human and agent judgments, and evaluate it with 241 human participants and 200 profile-conditioned LLM agents across four models. We find that nearest human--agent pairs in trait space achieve significantly higher TUX, suggesting that tacit alignment is structured by person-level characteristics rather than random similarity. Regression analyses show that TUX becomes more explainable as predictor sets become richer, with individual traits, decision-making styles, and confidence improving over aggregate trait-distance baselines. These findings suggest that tacit understanding between humans and LLMs is measurable, while revealing the limits of profile-based conditioning for capturing deeper representational alignment.
Delayed Momentum Aggregation: Communication-efficient Byzantine-robust Federated Learning with Partial Participation
arXiv:2509.02970v3 Announce Type: replace Abstract: Partial participation is essential for communication-efficient federated learning at scale, yet existing Byzantine-robust methods typically assume full client participation. In the partial participation setting, a majority of the sampled clients may be Byzantine, once Byzantine clients dominate, existing methods break down immediately. We introduce delayed momentum aggregation, a principle where the central server aggregates cached momentum from non-sampled clients along with fresh momentum from sampled clients. This principle ensures Byzantine clients remain a minority from the server's perspective even when they dominate the sampled set. We instantiate this principle in our optimizer DeMoA. We analyze the convergence rate of DeMoA, showing that DeMoA is Byzantine-robust under partial participation. Experiments show that, with 20% Byzantine ratio and only 10% partial participation rate, DeMoA achieves the best accuracy even when existing methods fail empirically.
Cost-aware Stopping for Bayesian Optimization
arXiv:2507.12453v5 Announce Type: replace Abstract: In automated machine learning, scientific discovery, and other applications of Bayesian optimization, deciding when to stop evaluating expensive black-box functions in a cost-aware manner is an important but underexplored practical consideration. A natural performance metric for this purpose is the cost-adjusted simple regret, which explicitly captures the trade-off between solution quality and cumulative evaluation cost. Existing stopping rules for Bayesian optimization are either heuristic, or are theoretically grounded but designed to optimize simple regret without accounting for evaluation costs; as a result, they provide no guarantees against unnecessary evaluations when costs are high. We propose a principled cost-aware stopping rule for Bayesian optimization that adapts to varying evaluation costs without heuristic tuning. Our rule is grounded in a theoretical connection to state-of-the-art cost-aware acquisition functions, namely the Pandora's Box Gittins Index (PBGI) and log expected improvement per cost (LogEIPC). When paired with either acquisition function, we prove that the resulting policy satisfies a theoretical guarantee bounding the expected cost-adjusted simple regret. Across synthetic tasks and empirical benchmarks including hyperparameter optimization and neural architecture size search, pairing our stopping rule with PBGI or LogEIPC usually matches or outperforms other acquisition-function--stopping-rule pairs in terms of cost-adjusted simple regret.
Dynamical local Fr\'echet curve regression in manifolds
arXiv:2505.05168v3 Announce Type: replace-cross Abstract: Under mild conditions, this paper derives a least-squares local linear Fr\'echet curve predictor for response and regressor evaluated in a separable Hilbert space. We obtain the conditions allowing the implementation of this local linear Fr\'echet functional predictor in the ambient L^{2}-space of vector functions, with values in the time-varying tangent space on a compact Riemannian manifold. An intrinsic local linear Fr\'echet curve predictor evaluated in such a manifold is secondly proposed, based on a weighted Fr\'echet mean approach. Its asymptotical optimality is proved. The simulation study and real-data application analyze the finite-sample performance of the empirical versions of both predictors, compared with a geodesic Nadaraya-Watson-type curve predictor. In the real-data application, the functional prediction of the time-varying spherical coordinates of the Earth's magnetic field is addressed, from the observation of the geocentric latitude and longitude of the satellite NASA's MAGSAT spacecraft.
Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning
arXiv:2605.31172v1 Announce Type: new Abstract: This work studies the convergence of two-timescale stochastic approximations (SA), a class of iterative algorithms that update two sets of parameters in fast and slow timescales respectively. Notable examples of two-timescale SA in reinforcement learning (RL) include temporal difference learning with gradient correction (TDC) and actor-critic methods. Previously, the stability (i.e., boundedness) and convergence of two-timescale SA were only established under i.i.d. noise. This work instead establishes the stability and convergence of two-timescale SA under Markovian noise, a setup that is more realistic in RL. Notably, we do not need to use any projection operator and the noise does not need to live in a compact space. Our key technical novelty is to control the fast timescale parameter with the running max of the slow timescale parameter, instead of with the current slow timescale parameter, as most prior works do. As a key application, we establish the first almost sure convergence of TDC with eligibility traces under off-policy learning with linear function approximation.
Deep-learning-based low-energy trigger algorithms for the Hyper-Kamiokande experiment
arXiv:2605.31391v1 Announce Type: new Abstract: Modern machine learning techniques have become increasingly important in particle physics because of their powerful pattern-recognition capabilities, including in real-time data acquisition where stringent runtime constraints apply. This paper details the performance of deep-learning-based trigger algorithms for a large water Cherenkov detector such as Hyper-Kamiokande aimed at low-energy neutrino events (below 7 MeV). The performance of custom neural-network supervised classifiers is shown alongside two anomaly-detection approaches trained solely on detector noise: a pure autoencoder and an energy-based model based on Manifold Projection--Diffusion Recovery (MPDR). The supervised model shows signal identification efficiencies of 76.7% for single electrons of 3 MeV kinetic energy, significantly exceeding signal efficiencies obtained from a traditional hit-count-based trigger of 26.4%, as does the MPDR approach with 31.8%. Runtime evaluations on GPU yield per-window inference latencies well below the millisecond scale, indicating that real-time operation is feasible.
Observation of resonant doublet and variable finesse in a tabletop meter-scale linear three-mirror cavity
arXiv:2505.03416v2 Announce Type: replace Abstract: Fabry-Perot cavities are widely used in current gravitational-wave detectors. In particular, they play a key role in frequency-dependent squeezing systems, enabling broadband quantum noise reduction. However, their ability to precisely control squeezing properties may be insufficient for the next generation of detectors. In this context, theoretical studies on linear three-mirror cavities have revealed promising features, such as resonance peak splitting and their equivalence to two-mirror cavities with variable finesse. In this paper, we report experimental observations of both the resonant doublet and variable finesse using a meter-scale implementation of a linear three-mirror cavity.
Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding
arXiv:2602.06161v2 Announce Type: replace Abstract: Parallel diffusion decoding can accelerate diffusion language model inference by unmasking multiple tokens per step, but aggressive parallelism often harms quality. Revocable decoding mitigates this by rechecking earlier tokens, yet we observe that existing verification schemes frequently trigger flip-flop oscillations, where tokens are remasked and later restored unchanged. This behaviour slows inference in two ways: remasking verified positions weakens the conditioning context for parallel drafting, and repeated remask cycles consume the revision budget with little net progress. We propose COVER (Cache Override Verification for Efficient Revision), which performs leave-one-out verification and stable drafting within a single forward pass. COVER constructs two attention views via KV cache override: selected seeds are masked for verification, while their cached key value states are injected for all other queries to preserve contextual information, with a closed form diagonal correction preventing self leakage at the seed positions. COVER further prioritises seeds using a stability aware score that balances uncertainty, downstream influence, and cache drift, and it adapts the number of verified seeds per step. Across benchmarks, COVER markedly reduces unnecessary revisions and yields faster decoding while preserving output quality.
Parity non-conservation in isotope chain of tin
arXiv:2605.26378v2 Announce Type: replace Abstract: We calculate parity non-conservation (PNC) amplitudes for all magnetic-dipole (M1) transitions within the ground $5p^2$ configuration of Sn. Among the transitions considered, the $^1$S$_0$-$^3$P$_1$ transition has the largest PNC amplitude and appears to be the most promising candidate for experiment. We also discuss a measurement method capable of achieving unprecedentedly high precision in a measurement of PNC in this transition. We argue that the most robust test should be based on ratios of PNC amplitudes for different isotopes, since the atomic-structure factor largely cancels in such ratios. We study the effect of the neutron skin on these isotope ratios using available nuclear data for Sn and show that the uncertainty associated with the neutron skin can be reduced to the $10^{-3}$ level relative to the isotopic variation of the PNC effect. Our results indicate that PNC measurements along a chain of Sn isotopes offer a realistic and sensitive probe of new physics.
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
arXiv:2605.30648v1 Announce Type: new Abstract: Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine learning tasks. We generalize this assumption to objectives whose curvature is an affine function of the objective value. This property is satisfied by a broad class of problems, including logistic regression, generalized linear models with a logistic link function, softmax policy gradient in reinforcement learning, and a class of neural networks. Under this assumption and gradient domination conditions, we establish a general convergence rate for the steepest descent method, and deterministic, diagonal variants of RMSProp and Adam. Our results imply that for logistic regression on separable data and the softmax policy gradient objective, sign GD converges linearly and is provably faster than GD. Furthermore, we show that for a class of two-layer neural networks on separable data, RMSProp and Adam can converge at a linear rate with a constant step-size and momentum parameter. Finally, we present a lower bound demonstrating that, under our assumption, RMSProp and Adam are provably faster than AdaGrad, AMSGrad, gradient descent, and heavy-ball momentum.
Diffusion Models Preferentially Memorize Prototypical Examples or: Why Does My Diffusion Model Love Slop?
arXiv:2605.30642v1 Announce Type: new Abstract: Generative models have a persistent limitation: their tendency to memorize training data can create legal liabilities and erode creative diversity. Understanding which samples are memorized in whole or in part, and under what conditions, therefore remains an important open problem. Here we answer the question "Are atypical or rare samples memorized first?" in the negative. We train diffusion models on strings generated according to the production rules of the Random Hierarchy Model (RHM), and find that samples composed of common substrings are preferentially memorized. This holds true even if the training data consists of entirely unique samples, indicating that deduplication at the data point level does not provide a meaningful privacy guarantee. Correspondingly we predict, then observe, delayed memorization for fat-tailed datasets (i.e., those with more atypical samples). This effect is amplified when fat-tails are introduced into high-level production rules. These together suggest that dataset diversity, particularly at higher levels of abstraction, plays an important role in staving off memorization. Finally, we identify an intermediate regime of partial memorization in which common substrings are learned first and subsequently overproduced during generation. If training is stopped in this regime, models will exhibit the reversion-to-the-mean blandness often derided as "slop".
Private Noise and Public Error in Collective Information Acquisition
arXiv:2605.30522v1 Announce Type: new Abstract: Collective information acquisition requires groups to combine personal evidence with social information while remaining coupled to the external state. Communication noise can affect this process, but the role of noise remains unclear. In an online experiment, 600 participants worked in four-person human groups estimating a room temperature across 25 rounds while receiving either faithful social information, comprehension noise in which each receiver saw independently perturbed social information, or production noise in which perturbations were stored before display and could be seen by multiple receivers. The thermometer cue was objectively veridical, but its reliability was subjectively uncertain and the unitless 50--250 room-temperature range created a task-induced conflict between displayed evidence and everyday temperature expectations. Production-noise groups spent more rounds tightly clustered around a wrong value than comprehension-noise groups (\(p=0.016\), group-level permutation). Production noise more often created a wrong common signal (\(p=0.025\), Fisher's exact test) and made that signal persist across more rounds (\(p=0.004\), permutation). Dynamic update models showed that production noise was not more harmful because people followed peers more strongly, but because the same peer influence acted on more correlated production-noise perturbations. Exploratory human analyses linked the mechanism to psychological patterns while a GPT-agent experiment clarified a boundary condition: GPT agents registered uncertainty through reduced confidence without reproducing human-scale production-noise vulnerability. Overall, noise did not simply degrade collective information acquisition. Comprehension noise could sometimes improve correction relative to the faithful control, whereas production noise could turn perturbations into common evidence and stabilize consensus on error.
Sequence models reveal diagnosis accumulation pathways beyond comorbidity burden in population-scale hospital data
arXiv:2605.30962v1 Announce Type: new Abstract: Aging trajectories vary among individuals of similar age and disease burden. Comorbidity indices, e.g. the Elixhauser index, summarize conditions cross-sectionally, but discard the timing, sequence, and pace of morbidity accumulation. Here we ask whether longitudinal hospital diagnosis histories contain information beyond age, sex, and comorbidity burden, and where it is concentrated. Using 13 years of Austrian inpatient data covering 7.4 million patients, we trained a visit-level contrastive transformer to encode diagnosis sequences and inter-admission timing into patient-history embeddings. In a downstream cohort of 1.7 million individuals, embeddings improved prediction over the Elixhauser-based comorbidity model for 93 of 131 incident ICD-10 disease-block outcomes, with a modest median AUC gain of 0.006. Gains concentrated in mental, musculoskeletal, nervous system, and metabolic disorders. We then evaluated event-free survival, defined as remaining alive without accumulating a second unrecorded ICD-10 disease block. The embedding model achieved an AUC of 0.726 versus 0.722 for the comorbidity model. However, among patients with similar age, sex, and comorbidity-model risk, those assigned high residual risk had 132--183 fewer event-free days over five years and observed event rates comparable to low-residual-risk patients more than a decade older. Together, these findings link the embedding's signal to the breadth, recency, and pace of prior disease accumulation.
Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders
arXiv:2602.10388v4 Announce Type: replace Abstract: The diversity of post-training data is critical for effective downstream performance in large language models (LLMs). Many existing approaches to constructing post-training data quantify diversity using text-based metrics that capture linguistic variation, but such metrics provide only weak signals for the task-relevant features that determine downstream performance. In this work, we introduce Feature Activation Coverage (FAC) which measures data diversity in an interpretable feature space. Building upon this metric, we further propose a diversity-driven data synthesis framework, named FAC Synthesis, that first uses a sparse autoencoder to identify missing features from a seed dataset, and then generates synthetic samples that explicitly reflect these features. Experiments show that our approach consistently improves both data diversity and downstream performance on various tasks, including instruction following, toxicity detection, reward modeling, and behavior steering. Interestingly, we identify a shared, interpretable feature space across model families (i.e., LLaMA, Mistral, and Qwen), enabling cross-model knowledge transfer. Our work provides a solid and practical methodology for exploring data-centric optimization of LLMs.
FoRA: Fisher-orthogonal Rank Adaptation for Parameter-Efficient Fine-Tuning
arXiv:2605.29317v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning(PEFT) has largely focused on LoRA and its accuracy-oriented variants, leaving the original goal of reducing trainable parameters has receivedcomparatively little attention. We introduce FoRA, which revisits this goal by reducing the number of adapted layers rather than adapter rank. FoRA selects task-informative layers via a single-pass diagonal Fisher score (under 1% of training cost) and trains the LoRA down-projection at selected layers on the Stiefel manifold, preserving column orthonormality and effective rank. FoRA consistently outperforms LoRA and DoRA at half their parameter budget, and falls within 0.7-0.8 accuracy points of AdaLoRA at one-quarter its parameter count, across five LLaMA-family backbones. Cross-architecture experiments on twelve backbones from the LLaMA, Qwen3, and Gemma families confirm consistent gains from 270M to 32B parameters. The two components combine super-additively: Fisher selection alone matches rank reduction at the same budget, while the Stiefel constraint provides the decisive additional gain.
Parameter-free Dynamic Regret: Time-varying Movement Costs, Delayed Feedback, and Memory
arXiv:2602.06902v2 Announce Type: replace Abstract: In this paper, we study dynamic regret in unconstrained online convex optimization (OCO) with movement costs. Specifically, we generalize the standard setting by allowing the movement cost coefficients $\lambda_t$ to vary arbitrarily over time. Our main contribution is a novel algorithm that establishes the first comparator-adaptive dynamic regret bound for this setting, guaranteeing $\widetilde{\mathcal{O}}(\sqrt{(M^2+MP_T)(T+\sum_t \lambda_t)})$ regret, where $P_T$ is the path length of the comparator sequence over $T$ rounds and $M$ is the maximal comparator norm. Our result recovers the optimal adaptive rates for both static and dynamic regret in OCO as the special case where $\lambda_t=0$ for all rounds. To demonstrate the versatility of our results, we consider two applications: OCO with delayed feedback and OCO with time-varying memory. We show that both problems can be translated into time-varying movement costs, establishing a novel reduction specifically for the delayed feedback setting that is of independent interest. A crucial observation is that the first-order dependence on movement costs in our regret bound plays a key role in enabling optimal comparator-adaptive dynamic regret guarantees in both settings.
Unfolding Generative Flows with Koopman Operators: Trajectory-Preserving Linearization
arXiv:2506.22304v3 Announce Type: replace Abstract: Continuous Normalizing Flows (CNFs) enable elegant generative modeling but remain bottlenecked by their iterative nature requiring costly sampling and lacking interpretability of the intermediate states. Recent approaches accelerate sampling by straightening trajectories or distilling endpoints, yet they treat the original generative process as a black box, discarding the teacher's intermediate dynamics. We propose a fundamentally different perspective: globally linearizing flow dynamics via Koopman theory to achieve trajectory-preserving linearization. By lifting a pre-trained Conditional Flow Matching (CFM) model into a higher-dimensional Koopman space, we represent its evolution with a single linear operator. Crucially, unlike boundary-only distillation, our method enforces infinitesimal consistency with the teacher's vector field along the full generative path. We derive a practical, simulation-free training objective that ensures this global alignment and yields two key benefits. First, sampling becomes one-step and parallelizable. Second, because the linearization is faithful to the dynamics, the Koopman operator provides unique insights on the generation. We demonstrate that this structure enables novel applications unavailable in prior approaches, including discovery of semantically coherent editing directions, inversion with a teacher-aligned linear operator and class-conditional spectral signatures. Empirically, our approach achieves competitive sample quality, while enabling spectral analysis and control of the entire trajectories of generative flows.
Tensor gradient flow for rod-like liquid crystals from molecular model with closure approximation by quasi-entropy
arXiv:2605.30735v1 Announce Type: cross Abstract: In tensor dynamics for liquid crystals derived from molecular models, a common problem is closure approximation. For rod-like molecules, the Bingham closure has proved to outperform other methods because it inherits the gradient flow structure of the molecular model, but is difficult to achieve efficient computations maintaining the gradient flow structure. We propose a closure approximation by the quasi-entropy that has been successfully applied to the free energy, based on which we construct the tensor gradient flow. The quasi-entropy closure has the same symmetry properties as the Bingham closure. The resulting tensor gradient flow is able to constrain the eigenvalues of the tensor within the physical range, guaranteeing the positive definiteness of the dissipation operator given by the higher-order tensors. The quasi-entropy closure is easy to implement since it can be reduced to minimizing an elementary function of three variables. As a result, we construct a numerical scheme preserving the eigenvalue constraints and energy dissipation, with the closure approximation decoupled from solving the scheme. Numerical simulations are carried out for the interface between the isotropic and the uniaxial nematic phase, as well as the defect evolutions, where the higher-order tensors indeed make a difference.
LIMCA: LLM for Automating Analog In-Memory Computing Architecture Design Exploration
arXiv:2503.13301v2 Announce Type: replace Abstract: Resistive crossbars enabling analog In-Memory Computing (IMC) have emerged as a promising architecture for Deep Neural Network (DNN) acceleration, offering high memory bandwidth and in-situ computation. However, the manual, knowledge-intensive design process and the lack of high-quality circuit netlists have significantly constrained design space exploration and optimization to behavioral system-level tools. In this work, we introduce LIMCA, a novel fine-tune-free Large Language Model (LLM)-driven framework for automating the design and evaluation of IMC crossbar architectures. Unlike traditional approaches, LIMCA employs a No-Human-In-Loop (NHIL) automated pipeline to generate and validate circuit netlists for SPICE simulations, eliminating manual intervention. LIMCA systematically explores the IMC design space by leveraging a structured dataset and LLM-based performance evaluation. Our experimental results on MNIST classification demonstrate that LIMCA successfully generates crossbar designs achieving $\geq$96% accuracy while maintaining a power consumption $\leq$3W, making this the first work in LLM-assisted IMC design space exploration. Compared to existing frameworks, LIMCA provides an automated, scalable, and hardware-aware solution, reducing design exploration time while ensuring user-constrained performance trade-offs.
Organizational Adaptation to Generative AI in Cybersecurity
arXiv:2506.12060v2 Announce Type: replace Abstract: Cybersecurity organizations are adapting to GenAI integration through modified frameworks and hybrid operational processes, with success influenced by existing security maturity, regulatory requirements, and investments in human capital and infrastructure. This qualitative research employs systematic document analysis and comparative case study methodology to examine how 25 studies from 2022 to 2025 document organizational adaptation of threat modeling frameworks, revealing a shift away from traditional signature-based systems toward AI-capable frameworks across three primary patterns: LLM integration for security applications, GenAI frameworks for risk detection and response automation, and AI/ML integration for threat hunting and matching. Organizations with mature infrastructures, particularly in finance and critical infrastructure, demonstrate higher readiness through structured governance, dedicated AI teams, and robust incident response processes, with central banks and financial institutions leading adaptation efforts under regulatory pressure. Successful integration requires human oversight of automated systems, attention to data quality and explainability, and sector-specific governance, though ongoing difficulties with privacy protection, bias reduction, personnel training, and adversarial defense persist. Notable imbalances between offensive and defensive GenAI capabilities create strategic concerns for security planning. The findings offer actionable insights for cybersecurity professionals and underscore the need for adaptive approaches, ethical frameworks, and staff development when managing AI-enhanced threats.
Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces
arXiv:2602.07864v2 Announce Type: replace Abstract: Spatial intelligence is crucial for vision--language models (VLMs), yet many scene-centric benchmarks evaluate unconstrained environments where a single image may admit multiple plausible 3D interpretations. We introduce SSI-Bench, a VQA benchmark for Structure-Centric Spatial Reasoning (SCSR) in constraint-governed spaces. Built from complex real-world 3D structures, it uses structural constraints from geometry, topology, and physical feasibility to make component relations more determinate from visual evidence. The benchmark contains 1,000 ranking questions spanning geometric and topological reasoning, where correct ordering requires resolving all candidate-wise 3D relations, imposing stronger demands on spatial understanding. It is created through a fully human-centered pipeline with over 400 researcher-hours of image curation, component annotation, and question design. Evaluating 31 VLMs reveals a large gap to humans: the best open-source model achieves 22.2% accuracy and the strongest closed-source model reaches 33.6%, while humans score 91.6%. Further results show that chain-of-thought reasoning brings only marginal gains, and error analysis reveals fundamental limitations in current models' spatial understanding within constraint-governed spaces. Project page: https://ssi-bench.github.io.
Neurodiversity in Agile Teams: Obstacles and Inclusion Barriers
arXiv:2605.30555v1 Announce Type: new Abstract: Context: Neurodiversity is increasingly recognized as a valuable dimension of workplace diversity. However, in agile software development teams, the interplay between teamwork practices and the inclusion of neurodivergent employees remains underexplored. Objective: The study aims to explore how teamwork quality in agile software development is currently practiced and discussed in the context of neurodiversity, and to identify organizational barriers that hinder the effective inclusion of neurodivergent developers. Method: We applied a mixed-method approach combining a web content analysis covering Reddit and LinkedIn with 11 semi-structured expert interviews from a corporate neurodiversity network in a German organization. Results: The analysis shows that teamwork practices are highly fragmented and shaped by individual adaptation rather than a shared standard. While agile practices and supportive tools can enable neurodivergent participation, rigid structures, stereotypes, and one-size-fits-all approaches often undermine inclusion. Organizational awareness and tailored adjustments remain insufficient. Conclusion: Agile practices can promote inclusive teamwork, yet their benefits are constrained by rigid organizational structures and limited awareness of neurodiversity. Harnessing neurodiverse strengths demands flexible organizational conditions and tailored support.