Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Design optimization and robustness analysis of rigid-link flapping mechanisms
arXiv:2503.21204v3 Announce Type: replace Abstract: Rigid link flapping mechanisms remain the most practical choice for flapping wing micro-aerial vehicles (MAVs) to carry useful payloads and onboard batteries for free flight due to their long-term durability and reliability. However, MAVs with these mechanisms require significant weight reduction to achieve high agility and maneuverability. One approach involves using single-DOF planar rigid linkages, which are rarely optimized dimensionally for high lift and low power, considering their sweeping kinematics and the unsteady aerodynamic effects. We integrated a mechanism simulator based on a quasistatic nonlinear finite element method with an unsteady vortex lattice method-based aerodynamic analysis tool within an optimization routine. We optimized three different mechanism topologies from the literature. Significant power savings were observed up to 34% in some cases, due to increased amplitude and higher lift coefficients resulting from optimized asymmetric sweeping velocity profiles. We also conducted a robustness analysis to quantify performance sensitivity to manufacturing tolerances. It provided a trade-off between performance and reliability and revealed the need for tight manufacturing tolerances and careful material selection. Finally, the analysis helped select the best mechanism topology, as we observed significant variation in sensitivity to manufacturing tolerances and peak input torque values across different topologies for a given design lift value. The presented unified computational tool can find application in flapping mechanism topology optimization, as it can simulate any generic single-DOF planar rigid linkage without supplying kinematics manually.
Psychological Safety Framework in Pull-based Open Source Projects
arXiv:2504.17510v4 Announce Type: replace Abstract: Psychological safety refers to the belief that team members can speak up, ask questions, and make mistakes without fear of negative consequences. Although psychological safety has been studied in traditional software teams, less is known about how it may appear in pull-based open-source software development, where contributors are self-directed and often collaborate voluntarily. This paper introduces a theory-informed framework for understanding how psychological safety may be reflected in pull request interactions. Drawing on psychological safety theory and prior work on software teams and open-source collaboration, the framework identifies observable interaction patterns related to feedback exchange, active participation, asking for input, and visible engagement from relevant project actors. To examine the framework empirically, we operationalize these patterns using nine observable variables from 60,684 pull requests across 26 popular GitHub repositories. The empirical results refine the framework by showing that visible engagement from contributors, reviewers, integrators, and other project members is positively associated with sustained participation, while interaction appears most useful when there is enough discussion without becoming excessive.
Unveiling Large Language Model Supply Chain: Structure, Domain, and Vulnerabilities
arXiv:2504.20763v2 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized artificial intelligence (AI), driving breakthroughs in natural language understanding, text generation, and autonomous systems. However, the rapid growth of LLMs presents significant challenges in the security and reliability of the Large Language Model Supply Chain (LLMSC), a complex network of open-source components, libraries, and tools essential for LLM development and deployment. Despite its critical importance, the LLMSC remains underexplored, particularly regarding its structural characteristics, domain composition, and security vulnerabilities. To address this gap, we conduct the first empirical study of the LLMSC, analyzing a curated dataset of open-source packages from PyPI and NPM across 14 functional domains. We construct a directed dependency graph comprising 13,486 nodes, 28,704 edges, and 180 unique vulnerabilities to investigate the structural characteristics of the LLMSC and analyze how security risks propagate through its dependency network. Our findings reveal that the LLMSC exhibits a locally dense, globally sparse topology, with 72.38% of dependency trees containing fewer than 5 nodes, while a few large trees dominate the ecosystem, accounting for 77.66% of all nodes. The graph is characterized by high-degree hubs, with the top 5 most connected nodes averaging 1,207 dependents each. Security analysis shows that critical vulnerabilities propagate to an average of 142.1 nodes at the second layer of dependency trees and peak at 237.8 affected nodes at the third layer. Notably, cascading risks are concentrated in critical hub nodes such as \texttt{transformers}, which directly or indirectly affect over 1,300 downstream packages. These findings provide quantitative insights into the structural and security dynamics of the LLMSC and emphasize the need for targeted mitigation strategies to enhance ecosystem resilience.
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
arXiv:2607.06008v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with external environments. However, most existing benchmarks implicitly assume a monolingual setting, where the entire execution process, including reasoning, tool invocation, and output generation, is conducted within a single language. In contrast, real-world applications often involve multilingual inputs and outputs within a unified workflow, yet the interaction between multilinguality and agentic execution remains underexplored. In this work, we introduce PolyWorkBench, a benchmark for evaluating LLM agents on multilingual long-horizon workplace workflows. PolyWorkBench consists of 67 tasks across five domains, including commerce, knowledge work, legal analysis, localization, and manufacturing, where agents must process heterogeneous multilingual inputs, perform iterative reasoning, invoke external tools, and produce structured outputs. To enable comprehensive evaluation, we propose a hybrid framework that combines structural grading, executable verification, and LLM-based semantic assessment. This design allows us to capture both functional correctness and linguistic consistency across complex workflows. Empirical results show that state-of-the-art LLM agents suffer significant performance degradation in multilingual workflow settings compared to monolingual counterparts. Our analysis suggests that multilinguality introduces compounding effects across reasoning and execution steps, highlighting the importance of jointly modeling language variation and procedural decision-making in agent evaluation.
Measuring the practice of shared-decision making (OPTION12): An Investigation into Open-sourced Smaller LLMs (OS-sLLMs) for Better Privacy and Sustainability
arXiv:2607.06127v2 Announce Type: replace Abstract: We present LLM4SDM, the first study of open-source smaller language models (OS-sLLMs) for automated assessment of shared decision making (SDM) using the Observer OPTION12 framework. Unlike previous work that relies on large commercial models and the shorter OPTION5 instrument, our study focuses on privacy-preserving locally deployable models and Dutch melanoma consultation transcripts. Using expert-annotated clinical consultations, we evaluate three general-domain and two medical-domain OS-sLLMs during a development-phase pilot study. Results show that general-domain models outperform medical-domain models, which exhibit substantial hallucination and instruction-following failures. Gemma3:12b achieves the strongest agreement with human annotations (Pearson r=0.51, Spearman \r{ho}=0.59). Item-level and qualitative analyses reveal systematic challenges related to temporal discourse reasoning, conversational role attribution, and evidence grounding. We further introduce a Judge-LLM consensus framework designed to support disagreement resolution among multiple models. Our findings suggest that while current OS-sLLMs cannot replace human annotators, they offer a promising foundation for privacy-preserving human-in-the-loop SDM assessment.
ToDMA: Large Model-Driven Massive Token Communications for Semantic Multiple Access
arXiv:2505.10946v3 Announce Type: replace Abstract: Token communications (TokenCom) is an emerging generative semantic communication paradigm, where tokens serve as compact representation units across modalities. Their contextual dependencies can be exploited by pretrained large models for semantic recovery. In this paper, we propose token-domain multiple access (ToDMA), a large-model-driven semantic multiple access scheme for massive token communications. ToDMA integrates unsourced random access with context-aware token processing. It enables massive uncoordinated devices to transmit tokenized source representations over common uplink resources. Specifically, each token index is associated with a shared modulation codeword, exposing token-level structure to the receiver for context-aware recovery. At the receiver, compressed sensing is first employed to jointly detect active tokens and estimate their corresponding channel state information (CSI) from the superposed signals. The source token sequences are then reconstructed by exploiting the consistency of token-associated CSI across multiple token positions. In the presence of token collisions, some active tokens may remain unassigned, leading to missing entries in the reconstructed token sequences. To recover these tokens, candidate-restricted masked-token prediction is performed using pretrained contextual models, thereby leveraging token-level context to mitigate collision effects. Simulation results on both image and text transmission tasks demonstrate that ToDMA reduces access latency while maintaining favorable token recovery and semantic reconstruction quality, showing its scalability for semantic multiple access.
Plasma-state metasurfaces for ultra-intensive field manipulation
arXiv:2505.15567v2 Announce Type: replace Abstract: High-power lasers offer ultrahigh intensities for plasma interactions, but they lack advanced techniques to control the properties of the fields, because no optical elements could withstand their high intensities. The vibrant field of metasurfaces has transformed modern optics by enabling unprecedented control over light at subwavelength through deliberate design. However, metasurfaces have traditionally been limited to solid-state materials and low light intensities. Extending the sophisticated capabilities of metasurfaces from solids into the plasma realm would open new horizons for high-field science. Here, we experimentally demonstrate plasma-state metasurfaces (PSMs) through the photonic spin Hall effect and stable-propagating vortex beam generation irradiated by intense light. Time-resolved pump-probe measurements reveal that the functionality of PSMs can persist for several picoseconds, making them suitable for controlling ultra-intense femtosecond lasers, even in state-of-the-art multi-petawatt systems. Harnessing the powerful toolkit of metasurfaces, this approach holds the promise to revolutionize our ability to manipulate the amplitude, phase, polarization, and wavefront of high-power lasers during their pulse duration. It also opens new possibilities for innovative applications in laser-plasma interactions such as compact particle acceleration and novel radiation sources.
Peer-Predictive Self-Training for Language Model Reasoning
arXiv:2604.13356v3 Announce Type: replace Abstract: Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predictive Self-Training (PST), a label-free fine-tuning framework in which multiple language models improve collaboratively by using a cross-model aggregate response as an internal training signal. Given a prompt, models generate responses sequentially; the final aggregated answer, which is often more reliable than individual responses in practice, serves as an internal reference for learning. We measure how informative each intermediate response is about the aggregate using pointwise mutual information (PMI), and use this signal to scale self-training updates: responses already aligned with the aggregate receive smaller updates, while less informative or misaligned responses receive larger ones. On mathematical reasoning benchmarks, including SimulEq, MATH-500-Numeric, and MultiArith, PST improves exact-match accuracy by 2.2--4.3 percentage points across Gemma-2-2B, LLaMA-3.2-1B, and Qwen2.5-1.5B, and reduces the average generator--verifier gap (GV-Gap) by 26--40%, while requiring no external supervision, no teacher--student hierarchy, and only cross-model interactions. These results suggest that peer-predictive feedback from cross-model generations can provide an effective mechanism for self-supervised language-model improvement.
PLURAL: A Global Dataset for Value Alignment
arXiv:2607.08034v1 Announce Type: new Abstract: Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse value systems. We introduce PLURAL, a large-scale, value-focused preference dataset grounded in the Integrated Values Survey (IVS), a nationally representative survey spanning 92 countries. Using a two-stage generation pipeline, we transform survey responses into synthetic preference triplets that preserve normative value signals while producing realistic scenarios. We release an initial version of PLURAL containing ~500,000 preference triplets representing people in 20 diverse countries. We evaluate PLURAL in three ways: (i) dataset-level validation showing that it preserves both cross-country value differences and within-country diversity from the original survey; (ii) automated evaluation showing that training on PLURAL improves alignment with target countries' cultural profiles, reducing mean absolute error by up to 27.7% relative to strong baselines; and (iii) blind human evaluation with 176 evaluators in India, Brazil, and Japan, who judge PLURAL-aligned responses as more representative of their national values. Together, these results show that PLURAL contains learnable signal for value steering, offering a scalable resource for pluralistic alignment. Dataset: https://huggingface.co/datasets/agdhruv/plural-alignment
On the Convergence of Belief Propagation for Multipath Data Association in Target Tracking
arXiv:2607.08521v1 Announce Type: new Abstract: Belief propagation (BP) is widely used for data association (DA) in target tracking. Existing convergence analyses of BP for DA address only the two-way correspondence between targets and measurements, where each target generates at most one measurement per scan. Multipath DA (MPDA) allows a single target to produce multiple measurements via distinct propagation paths, creating a three-way correspondence among targets, paths, and measurements, for which a complete convergence proof has not yet been provided. We provide such a proof for the BP updates in MPDA, establishing convergence to a unique fixed point. Simulations illustrate the convergence behavior of BP in MPDA and demonstrate a favorable accuracy--efficiency trade-off relative to both single-scan and two-scan variants of the multiple-detection multiple-hypothesis tracker.
An exact information theory of generalization phase transitions in Bayesian diffusion models
arXiv:2607.08041v1 Announce Type: new Abstract: How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model time-reverses diffusion by inferring which past training sample produced its current restricted observation using the Bayesian posterior. This model class generalizes existing analytical diffusion models that use spatially local information restriction. We show that spatially local BIRD models closely approximate trained diffusion models \textit{early in training}, across different architectures such as UNets and DiTs. Under minimal assumptions on the data distribution, we identify an information-theoretic phase boundary between memorization and generalization in the joint space of amount of training data, time in the reverse generative process, and amount of information restriction: a BIRD model memorizes when the mutual information between its restricted noisy observations and the training data exceeds the log number of training points, and it generalizes otherwise. Experiments across a range of datasets confirm our theoretically predicted location for the transition. We find that generation proceeds near the edge of memorization: both spatially local BIRD models and early-training diffusion models track the memorization-generalization phase boundary by increasingly restricting information over time. Overall, our results reveal a fundamental role for information restriction in generative AI to circumvent the curse of dimensionality.
Using covariance of node states to design early warning signals for network dynamics
arXiv:2505.15982v2 Announce Type: replace Abstract: Real-life systems often experience regime shifts. An early warning signal (EWS) is a quantity that attempts to anticipate such a regime shift. Because complex systems of practical interest showing regime shifts are often dynamics on networks, a research interest is to design EWSs for networks, including determining sentinel nodes that are useful for constructing high-quality EWSs. Previous work has shown that the sample variance is a viable EWS including in the case of networks. We explore the use of the sample covariance of two nodes, or sentinel node pairs, for improving EWSs for networks. We perform analytical calculations in four-node networks and numerical simulations in larger networks to find that the sample covariance and its combination over node pairs is inferior to the sample variance and its combination over nodes; the latter are previously proposed EWSs based on sentinel node selection. The present results support the predominant use of diagonal entries of the covariance matrix (i.e., variance) as opposed to off-diagonal entries in EWS construction.
Economised path integrals
arXiv:2607.06414v2 Announce Type: replace Abstract: The Hessian of the ring polymer spring potential in the standard Trotter path integral is a $P\times P$ symmetric circulant matrix with a centroid eigenvalue of zero. All such matrices commute and are diagonalised by the same bead to normal mode transformation matrix, and their eigenvalues contain $\lceil P/2\rceil-1$ degenerate pairs by symmetry. However, this still leaves some freedom to improve on the Trotter approximation: one can optimise the remaining $\lfloor P/2\rfloor$ independent non-zero normal mode frequencies to fit the exact quantum mechanical radii of gyration of harmonic ring polymers with frequencies in the range $0\le\omega\le\omega_{\rm max}$, where $\omega_{\rm max}$ is the maximum physical frequency in the problem of interest. The optimisation involves solving a simple least squares problem for the optimum (economised or "Eco") internal mode frequencies. The remainder of the calculation then proceeds in the same way as a Trotter path integral calculation. An example application to hexagonal ice shows that the convergence of the Eco path integral is comparable to that of the 4th order Suzuki-Chin path integral, but with purely 2nd order Trotter effort. There is no need to calculate the projected Hessians that arise in the Suzuki-Chin method by finite differences, there is no need to develop any new estimators for observables, and once the Eco frequencies have been calculated the implementation of the Eco path integral involves changing just a few lines of a Trotter path integral code. To provide a more impressive example we have implemented the Eco method in GPUMD and used it to converge the (negative) thermal expansion coefficient and the constant pressure heat capacity of MOF-5 with a machine-learned neuroevolution potential.
STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning
arXiv:2607.06629v2 Announce Type: replace Abstract: Brain age - the age inferred from a physiological recording - is an emerging biomarker whose deviation from chronological age tracks neurological and psychiatric burden, and EEG is an attractive substrate for it because it is cheap, portable, and temporally rich. Yet EEG brain-age models must contend with cross-site montage heterogeneity, small labelled cohorts, and dominant subject-level non-stationarity, and few EEG foundation models have been shown to deliver competitive age regression across the full pediatric to older adult range in which such a biomarker would actually be deployed. We introduce STST-JEPA, a self-supervised transformer for resting-state and task EEG, pretrained on 47,703 sessions spanning ages 5-81 from the brain.space and Healthy Brain Network (HBN) corpora. The model combines a latent-prediction objective - predicting masked-token representations against an EMA-of-tokenizer target - with an auxiliary signal-reconstruction term, applied to 30-second multi-channel windows under spatiotemporal block masks. A lightweight attentive probe trained on frozen pretrained embeddings achieves a best held-out-validation mean absolute error of 3.06 years (r = 0.924) for age regression on 3,367 sessions, against a predict-the-mean baseline of approximately 10 years MAE. With light task-specific finetuning of the model's final layers, the same pretrained encoder achieves rank-1 placements - with the model's native 30-second windows - on the public NeuralBench x brain.space EEG leaderboard for sex classification (balanced accuracy 0.911), age prediction (r = 0.749), and psychopathology composite regression (r = 0.215). We further show that the model's age-prediction residual is negatively correlated with cognitive efficiency over several tasks we examined.
It Takes Few to TANGO: A Quantized Distributed Model for Binaural Speech Enhancement
arXiv:2607.08645v1 Announce Type: new Abstract: Neural network-based multichannel speech enhancement systems achieve strong enhancement performance, but their computational and memory requirements limit deployment on resource-constrained devices. This paper investigates low-precision inference for TANGO, a hybrid distributed binaural speech enhancement system combining neural mask estimation with spatial filtering. We evaluate post-training quantization and quantization-aware training for the neural components, and analyze how quantization errors in the mask estimators propagate through the downstream spatial filtering stage. Our analysis shows that, although quantization degrades intermediate mask estimates, the spatial filtering stage compensates for most quantization-induced errors. Leveraging this robustness, we simplify TANGO into MN-TANGO, reducing both model size and computational complexity while maintaining comparable final performance. By combining INT8 weight-and-activation quantization with ERB compression and grouped recurrent layers, the most compact MN-TANGO reaches 4.65 MMAC/s and 0.177 MB.
Fair Document Valuation in LLM Summaries via Shapley Values
arXiv:2505.23842v5 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly power search engines and AI assistants that retrieve and summarize content from many sources. By serving answers directly, these systems obscure the original content creators' contributions, threatening the compensation that sustains a healthy content ecosystem. We frame this as a problem of fair document valuation and compensation, and propose a framework based on the Shapley value. Because exact Shapley computation is prohibitively expensive at scale, we develop Cluster Shapley, an approximation that groups semantically similar documents via LLM embeddings and computes Shapley values at the cluster level, with formal bounds on both the approximation error and the induced revenue-attribution error. On Amazon product review data, off-the-shelf approximations such as Monte Carlo sampling and Kernel SHAP perform suboptimally in LLM settings, whereas Cluster Shapley substantially improves the efficiency--accuracy frontier. Simple attribution heuristics (e.g., equal or relevance-based allocation), though computationally cheap, yield highly unfair outcomes. Our approach is agnostic to the exact LLM used, the summarization process used, and the evaluation procedure, which makes it broadly applicable to a variety of summarization settings.
Crystalline metal flakes: Platforms for advanced plasmonics and hybrid 2D material architectures
arXiv:2604.22988v2 Announce Type: replace Abstract: Crystalline noble metal flakes are emerging as versatile platforms in nanophotonics, enabling a broad range of optical phenomena and applications. Their atomically flat surfaces, high crystallinity, and superior optical quality open new avenues in advanced plasmonics, quantum light generation, and hybrid photonic systems. In contrast to conventional polycrystalline metal films, which typically suffer from higher optical losses due to grain boundaries, surface roughness, and structural disorder, these monocrystalline flakes provide minimal scattering and enhanced performance. They serve as templates for precise nanostructuring through techniques like focused-ion beam (FIB) milling and are crucial for advanced applications in sensing and optoelectronics. Additionally, they facilitate frontier research in quantum plasmonics, enabling fundamental studies of nonlocal optical effects and the generation of nonclassical light. Furthermore, the well-defined $\{111\}$ facets of these flakes host Tamm--Shockley surface states that support 2D plasmons coexisting with bulk modes. At near-infrared wavelengths and beyond, crystalline flakes act as nearly ideal metallic mirrors, featuring surface roughness limited only to atomic terrace steps, making them highly suitable for integration with 2D materials in hybrid photonic architectures. This review surveys the key roles these flakes play, highlighting recent developments and discussing future prospects while emphasizing their unique benefits in addressing fundamental and applied challenges in modern nanophotonics.
Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism
arXiv:2506.01260v3 Announce Type: replace Abstract: Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to communication bottlenecks. While existing compression techniques are effective in data-parallel, they do not extend to model parallelism. Unlike data-parallel training, where weight gradients are exchanged, model-parallel requires compressing activations and activation gradients as they propagate through layers, accumulating compression errors. We propose a novel compression algorithm that compresses both forward and backward passes, enabling up to 99% compression with no convergence degradation with negligible memory/compute overhead. By leveraging a recursive structure in transformer networks, we predefine a low-dimensional subspace to confine the activations and gradients, allowing full reconstruction in subsequent layers. Our method achieves up to 100x improvement in communication efficiency and enables training billion-parameter-scale models over low-end GPUs connected via consumer-grade internet speeds as low as 80Mbps, matching the convergence of centralized datacenter systems with 100Gbps connections with model parallel.
SimdQuickHeap: The QuickHeap Reconsidered
arXiv:2604.25681v2 Announce Type: replace Abstract: Priority queues are data structures that maintain a dynamic collection of elements and allow inserting new elements and removing the smallest element. The most widely known and used priority queue is likely the implicit binary heap, even though it has frequent cache misses and is hard to optimize using e.g. SIMD instructions. We introduce the SimdQuickHeap, a variant of the QuickHeap that was introduced by Navarro and Paredes in 2010. As suggested by the name, the data structure bears some similarity to QuickSort. We modify the data layout of the original QuickHeap to have all pivots adjacent in memory, with elements between consecutive pivots stored in dedicated buckets. This allows efficient SIMD implementations for both partitioning of buckets and scanning the list of pivots to find the bucket to append newly inserted elements to. The SimdQuickHeap has amortized expected complexity $O(\log n)$ per operation, which improves to $O(\frac 1W\log n)$ in non-degenerate cases, where $W$ is the number of words in a SIMD register. In this case, the I/O-complexity is amortized $O(\frac 1B)$ per push and $O(\frac 1B \log_2 \frac nM)$ per pop. In synthetic benchmarks, the SimdQuickHeap is $1.2\times$ to $1.7\times$ as fast as the monotone radix heap, the next-best competitor, and $1.4\times$ to $2.8\times$ as fast as the superscalar sample queue, the fastest comparison-based priority queue. The SimdQuickHeap needs around $1.5\log_2 n$ comparisons and $\log_2 n$ nanoseconds per pair of push and pop operations. On graph benchmarks with Dijkstra's shortest path algorithm and Jarn\'ik-Prim's minimum spanning tree algorithm, the SimdQuickHeap is consistently the fastest.
DeepTutor: Towards Agentic Personalized Tutoring
arXiv:2604.26962v3 Announce Type: replace Abstract: Education is one of the most promising real-world applications for Large Language Models (LLMs). However, current LLMs rely on static pre-training knowledge and lack adaptation to individual learners, while existing RAG systems fall short in delivering personalized, guided feedback. To bridge this gap, we present DeepTutor, a fully open-source agentic framework that unifies citation-grounded problem tutoring with difficulty-calibrated question generation. A hybrid personalization engine couples static knowledge grounding with dynamic learner memory, continuously adapting each interaction to the student's evolving needs. The same personalization substrate further extends to adaptive learning workflows, interactive books, and proactive multi-channel tutoring agents. To evaluate personalized tutoring, we introduce TutorBench, an interactive benchmark incorporating customized learner profiles grounded in university-level curricula across five domains. We further propose an LLM-based first-person interactive evaluation protocol that conducts assessments via a profile-driven student simulator. Complementary evaluations on established benchmarks, supported by human-alignment and ablation studies, confirm the framework's robustness and general utility. Results show that DeepTutor improves personalized metrics by 10.8\% on average and strengthens general agentic reasoning across five backbone models by 29.4\%.
A stochastic agent-based extension of the GSM2 model for particle therapy: cell-cycle dynamics, dose-rate dependence, and fractionation effects
arXiv:2604.27630v2 Announce Type: replace Abstract: Accurately linking microscopic energy deposition from ionizing radiation to emergent biological outcomes remains a central challenge in radiobiological modelling, particularly when stochastic damage induction, cell-cycle dynamics, and spatial organisation within irradiated tissues must be treated explicitly and consistently across scales. To address this, we introduce a stochastic agent-based radiobiological modelling framework for simulating biological response to particle irradiation, developed as an explicit single-cell extension of the Generalized Stochastic Microdosimetric Model (GSM2). Each cell is represented as an autonomous agent whose internal state, including DNA lesion counts, cell-cycle phase, and oxygenation level, evolves according to a continuous-time Markov chain driven by GSM2 transition rates. Radiation-induced damage induction, repair, misrepair, cell-cycle progression, proliferation, and migration are treated as competing stochastic events resolved through a next-event, event-driven algorithm, which provides computationally efficient scaling with system size while preserving full single-cell resolution. The framework is applied to three-dimensional tumour spheroids irradiated with 1H and 12C ions across a range of energies and dose rates. We characterise the spatiotemporal evolution of cell-cycle phase composition and spheroid volume following irradiation, and examine the dependence of cell survival on dose rate over four orders of magnitude. Several empirically established trends in biological response, including the dose-rate dependence of cell survival, its attenuation at high LET, and the inverse dose rate effect in split-dose irradiation, emerge from the model through the explicit coupling of particle arrivals, damage accumulation, and repair kinetics, without recourse to empirical correction factors as typically done.
Multiphysics embedding localized orthogonal decomposition for thermomechanical coupling problems
arXiv:2507.13644v2 Announce Type: replace Abstract: Multiscale thermomechanical problems in highly heterogeneous media are challenging because the elastic, thermal, and coupling coefficients may vary on unresolved spatial scales. We propose a multiphysics-embedding localized orthogonal decomposition (ME-LOD) method in which displacement and temperature correctors are generated by a coupled static operator. The corrector problems are localized to coarse-grid patches and solved in the kernel of a projective quasi-interpolation operator. We prove uniform inf-sup stability on the global fine-scale kernel and on all zero-extension patch kernels, establish exponential decay of the coupled correctors and the resulting multiscale basis functions, and derive spatial approximation and fully discrete reduction estimates. Numerical experiments demonstrate that, for the tested periodic, random, and high-contrast coefficient fields, ME-LOD attains smaller errors than the comparison method at the same coarse resolution and patch size and can reach a prescribed accuracy with fewer oversampling layers. Although each coupled local corrector is more expensive than a decoupled corrector, the improved localization yields a favorable overall accuracy-to-cost balance in the reported tests.
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
arXiv:2605.03378v2 Announce Type: replace Abstract: Large Language Model (LLM) agents are increasingly deployed as task-oriented software systems that use runtime context to decide and act on behalf of users. This delegation model makes prompt injection especially dangerous: an attacker can hide a context-aware instruction inside evidence the agent must use to decide what to do. Existing benchmarks and defenses largely miss this setting. Benchmarks often use context-insensitive tasks where the user prompt already specifies the intended action, together with generic attack payloads independent of context. Existing defenses also do not capture the causal support from runtime evidence to concrete actions, which makes them incomplete and ineffective for context-dependent tasks. We present AgentLure, a benchmark for context-dependent tasks under context-aware prompt injection. AgentLure spans four agentic domains and eight attack vectors across six attack surfaces. To defend this setting, we propose ARGUS, a causal-provenance auditor for LLM agents. Instead of relying only on tool authorization or suspicious-context detection, ARGUS verifies whether each proposed action has a complete benign causal justification. It builds an influence-provenance graph, labels runtime spans, grounds action arguments in supporting evidence, and releases an action only when benign evidence entails it and task invariants hold. On AgentLure, ARGUS reduces attack success rate from 28.8% to 3.8% while preserving 87.5% clean utility, significantly outperforming existing defenses in the security-utility tradeoff.
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite Comparison
arXiv:2605.03713v3 Announce Type: replace Abstract: Specialized accelerators dominate AI workloads, but CPUs remain critical for orchestrating accelerators and running daily services. CPU performance therefore shapes end-to-end system efficiency, making benchmarks reflect modern workloads and bottlenecks. Yet it remains unclear whether the newest general-purpose CPU benchmark suite changes the architectural conclusions drawn from prior SPEC CPU generations. We present the first comprehensive characterization of SPEC CPU2026 across nine recent Intel, AMD, Ampere, and Nvidia platforms. Compared with SPEC CPU2017, SPEC CPU2026 increases instruction volume and memory footprint and shifts pressure toward emerging bottlenecks, especially instruction-cache stress. These shifts raise two practical questions: how much of the new suite is needed to preserve behavioral coverage, and how should that coverage be interpreted relative to modern domain-specific suites? Using clustering-based representativeness analysis, we identify compact subsets of 4-5 workloads per group that preserve 96.4-99.9% of full-suite behavior. We then compare SPEC CPU2026 against SPEC CPU2017, DCPerf, and MLPerf using cross-suite microarchitectural metrics. SPEC CPU2026 remains a complementary general-purpose suite: it moves closer to datacenter-like frontend pressure than prior SPEC CPU generations, while remaining less vector-intensive than MLPerf and less frontend-extreme than DCPerf. Finally, case studies on page sizes and allocators, prefetching, compiler optimizations, ISA sensitivity, many-core scaling, and rolling round-robin proxy workloads show that SPEC CPU2026 supports architectural studies beyond aggregate scores. Overall, SPEC CPU2026 updates the standardized general-purpose CPU baseline for the next decade of architecture evaluation.
Inverse-designed release-free optomechanical crystal with high photon-phonon coupling
arXiv:2605.03910v2 Announce Type: replace Abstract: Interactions between light and mechanics provide a powerful interface between optical and microwave-frequency signals, with applications spanning classical signal processing and quantum technologies. High-performance optomechanical devices require both strong photon-phonon coupling and tolerance to parasitic laser heating. Release-free optomechanical crystals provide improved thermal anchoring compared to suspended nanobeams, but have so far exhibited weaker vacuum optomechanical coupling rates, leaving a trade-off between coupling strength and thermal robustness. Here, we largely close this gap: we design and experimentally demonstrate a release-free silicon optomechanical crystal with a record vacuum optomechanical coupling rate of about $g_\text{OM} / (2 \pi) = 800$ kHz, comparable to suspended state-of-the-art devices. The resulting optomechanical scattering rate $\Gamma_\text{OM}/(2 \pi)= 1.1$ kHz is nearly twice that of previous release-free implementations. This performance is achieved by combining physics-guided human intuition with a multiphysics inverse-design algorithm introduced here for resonant optomechanical structures. Beyond the specific device demonstrated, the inverse-design framework is applicable to co-optimizing optical and mechanical resonances and eigenmodes more broadly. These results strengthen release-free optomechanical crystals as a platform for fast, low-noise classical and quantum optomechanics.