arXiv:2607.17095v1 Announce Type: new
Abstract: Accurate short-term wind power forecasting is essential for grid stability and operational planning, yet remains challenging due to the complex interactions between atmospheric conditions and turbine dynamics. However, existing methods fail to effectively incorporate weather forecasting with wind turbine data (i.e., SCADA), leading to suboptimal solutions. To address this, we introduce a multimodal framework that integrates historical point-based SCADA data with grid-based Numerical Weather Prediction (NWP) forecasts, which is challenging due to heterogeneous input and the complex physical wind-turbine interactions. Our approach first explicitly decomposes inputs into scalar and vector features to better capture both site-specific and geometric dependencies and then incorporates a geometric encoder to extract rotation-invariant features from wind vectors. We further leverages a Fourier Neural Operator (FNO) architecture, which performs global convolutions in the frequency domain to efficiently model long-range spatiotemporal relationships. Extensive experiments on three real-world wind farms, with weather forecasting data, demonstrate that our model consistently outperforms state-of-the-art baselines, highlighting the effectiveness of its physically-informed design. The core implementation of our method is publicly available at: https://github.com/shawn-sypiao/GWPF.
Science Journals
arXiv:2607.17090v1 Announce Type: cross
Abstract: We study the expressivity of shallow polynomial neural networks (PNNs) with monomial activation functions over finite fields. For a given architecture, we define a neuromanifold as the image of the map from all possible network weights into the product of polynomial rings. We quantify the expressivity by the cardinality of the neuromanifold, and derive a natural lower and upper bound. This leads to counting rational points over finite fields, a problem closely linked to the Weil conjectures. Finally, we present an architecture that exhibits a striking difference in the neuromanifolds when considered over a characteristic zero versus a finite-characteristic field, illustrating the critical role of field characteristic in the notion of expressivity.
arXiv:2607.16367v1 Announce Type: new
Abstract: On flat superhydrophobic surfaces, droplet rebound is well described by a single inertio-capillary time scale, yielding a contact-time that is independent of impact energy. This single-mode response reflects the radial symmetry of flat-plate impacts. We demonstrate that an anisotropic geometric constraint, imposing a fixed spreading length along one axis, breaks this degeneracy and splits the rebound into a reciprocal pair of inertio-capillary modes. The fixed length also couples the contact-time to the Weber-dependent maximum spread, introducing an impact-energy dependence absent on the flat plate. We realize this constraint with grooved substrates, simulated using a non-ideal, entropic, multiple-relaxation-time lattice Boltzmann method and validated against the experiments of Chantelot et al. Extending their blob model from a single transverse scale to the reciprocal pair, we organize both modes through a geometric blob number and relate their time scales to the Weber number and groove width. We show that on non-wetting grooves the reciprocal modes are recovered directly, and explore the effects of finite wall affinity, using competition between the two modes to explain an observed two-branch structure in the contact-time response on mildly wetting, superhydrophobic grooves. Predictions tied to global energy balance reproduce cleanly across all conditions, while those tied to the details of the droplet's spread morphology are approximate but directionally correct. These results show that anisotropic confinement turns contact-time reduction from a question of accelerating a single rebound mode into one of selecting between conjugate inertio-capillary modes.
arXiv:2607.17042v1 Announce Type: new
Abstract: Humanoid robots have become increasingly popular in applications such as social interaction, education, and service roles, which drives the need for more natural and efficient human-robot interactions. However, currently available humanoid heads often face limitations, including high costs, mechanical complexity, and limited adaptability across diverse environments. To address these challenges, we present an articulated humanoid robot head designed for a receptionist role, integrating a mechanical structure with 21 degrees of freedom (DoF), including mechanisms for the mouth, eyes, eyebrows, and neck, and covered with realistic silicone skin to achieve a human-like appearance and expression. The system integrates a model-based architecture that combines SCRFD, ArcFace, and ByTetrack for face recognition and Llama and Whisper for natural language processing, with hardware support enabling real-time operations and human re-identification. The conversational ability and re-identification capabilities of the humanoid robot head were quantitatively measured, while its emotional expressiveness and human likeness were evaluated through a user study, achieving an average human likeness score of 4.13 out of 5.
arXiv:2607.17316v1 Announce Type: new
Abstract: The softmax policy $\pi(a \mid s) \propto \exp(\beta Q(s,a))$ is the default model of stochastic choice in reinforcement learning (RL). Various justifications based on robustness, exploration, and optimization have been offered in the RL literature, but none uniquely derives the softmax form from first principles. This leaves a basic tension unresolved: the entropy bonus in the soft Bellman equation violates the Independence axiom that underwrites the Markov decision process (MDP) reward structure. We dissolve this tension by distinguishing two kinds of randomness: chance and choice. By restricting von Neumann-Morgenstern (VNM) Independence to environmental lotteries over base prospects, we show that imposing independence of irrelevant alternatives (IIA) and monotonicity on the policy and value functions at choice nodes uniquely determines the Boltzmann policy, the entropy-regularized representation, and the soft Bellman equation. The choice between the soft and hard Bellman equations thus reduces to a design decision: whether the agent values its own ability to choose. We develop RL-specific consequences, including return monotonicity and convergence under generalized discounting, and synthesize the independent lines from economics and information theory that arrive at the same structure, offering a normative assessment of when IIA is appropriate for agent design.
arXiv:2607.17409v1 Announce Type: cross
Abstract: We study the problem of sequentially evaluating a new large language model (LLM) on a fixed question set using historical performance data from prior LLMs. Our goal is to construct a confidence sequence (CS) for the model's capability on this question set and to design active querying rules that shrink the CS width as quickly as possible. For CS construction, we invert a family of test supermartingales and focus on two representative approaches: a reverse information projection (RIPr)-based approach and a testing-by-betting-based approach. We first study these approaches under an oracle setting, and demonstrate the oracle optimality of the RIPr-based construction. We then propose a growth-oriented querying rule that aims to maximize the worst-case one-step expected log-increment over the endpoints of the current CS. In practice, we build these test supermartingales and the querying rule on predictions of question-level correctness learned from historical data. We then analyze the shrinkage behavior of the resulting CSs and identify two key factors that slow the shrinkage rate of CSs: accumulated prediction mismatch and the spikiness of the querying distribution. Finally, motivated by this analysis, we propose several mixture querying rules that combine growth-oriented querying, prediction refinement, and uniform exploration, trying to mitigate the effects that slow the shrinkage rate. We provide experiments comparing different querying rules for the RIPr-based and testing-by-betting-based CSs across several synthetic testing datasets. Interestingly, we observe that the simplest querying rule, uniform sampling, can sometimes outperform more adaptive querying rules for both methods.
arXiv:2607.17421v1 Announce Type: cross
Abstract: The strong chromatic index $\chi'_s(G)$ is the smallest number of colours needed to colour the edges of a graph $G$ so that any two edges at distance at most $2$ receive different colours. Using the \emph{local flag algebra} framework introduced in a companion paper, we prove $\chi'_s(G) \leq 1.73\,\Delta(G)^2$ for every graph $G$ of maximum degree $\Delta(G)$, $\chi'_s(G) \leq 1.6255\,\Delta(G)^2$ for every bipartite $G$, and $\chi'_s(G) \leq 1.6633\,\Delta_A(G)\,\Delta_B(G)$ for every bipartite $G$ of side maximum degrees $\Delta_A(G), \Delta_B(G)$ with rational $\Delta_B(G)/\Delta_A(G) \in (0, 1]$, provided $\Delta(G)$, $\Delta_A(G)$, $\Delta_B(G)$ are sufficiently large. These three bounds make progress towards three established conjectures: those of Erd\H{o}s-Ne\v{s}et\v{r}il (1985) for general graphs, Faudree-Gy\'arf\'as-Schelp-Tuza (1989) for bipartite graphs, and Brualdi-Quinn Massey (1993) in the asymmetric bipartite setting.
Additionally, for the random bipartite graph $G \sim G(n_A, n_B, p)$ at constant $p \in (0,1)$ and bounded aspect ratio $\max(n_A, n_B) = O(\min(n_A, n_B))$, we prove the Brualdi-Quinn Massey bound $\chi'_s(G) \leq \Delta_A(G)\,\Delta_B(G)$ asymptotically almost surely.
arXiv:2607.16534v1 Announce Type: new
Abstract: Reinforcement learning (RL) has achieved strong results in control, yet learned policies remain brittle to changes in dynamics, action spaces, observation spaces, or goals, a critical limitation for real-world deployment. Existing benchmarks offer limited diversity and complexity, making it difficult to rigorously study transfer, multi-task learning, and meta-learning in RL. We introduce Building2Building (B2B), a large-scale suite of realistic Heating, Ventilation, and Air Conditioning (HVAC) control environments built on EnergyPlus, a state-of-the-art building simulator. B2B is fully compatible with the Gymnasium interface and features a parametric building generator, enabling the systematic generation of diverse building configurations with heterogeneous observation and action spaces. Based on this suite, we define benchmark tasks targeting key open challenges in RL, including goal adaptation, dynamics adaptation, action-space shifts, and cross-domain transfer. By providing a large-scale, diverse, and physically grounded testbed with standardized evaluation protocols, B2B enables systematic investigation of generalization and transfer in continuous control. Beyond advancing research on generalization in RL, this new benchmark also carries significant societal implications by enabling improved HVAC control at scale, one of the most energy-intensive systems in buildings.
arXiv:2607.17447v1 Announce Type: cross
Abstract: Language models produce probabilities over words, but professional decisions require uncertainty over meaningful states such as diagnoses, hypotheses or operational conditions. A model's printed numerical confidence does not establish reliability. We introduce a semantic map: a prespecified, testable bridge from probabilities over verbal responses to probabilities over declared states, formulated as semiparametric inference for a finite-valued latent state. A reference model defines the target posterior, the language model supplies an unrestricted conditional distribution over verbal responses, and held-out calibration connects them. We derive posterior-error bounds and conditions for existence, uniqueness, stability and sequential Bayesian updating. Crucially, language probabilities depend on the prompt's lexical form, whereas the target posterior is unchanged by information-equivalent rewording. We test the method on professional market text compiled from Federal Reserve economic and financial series and on controlled simulations with exact posteriors. Across two fitted language models, language-derived probabilities outperform printed numerical probabilities, recover held-out posteriors with valid uncertainty coverage, remain largely stable under paraphrasing, and respond appropriately to altered evidence. The broader implication is that prompt engineering optimizes a wording-dependent response, whereas scientific and professional use requires validated stability of application-relevant meaning. The semantic map turns this general concern into a testable statistical problem and, when its acceptance conditions hold, yields an auditable posterior estimate. The same principle offers a template for auditing classifications, recommendations and other fluent responses that may conceal semantic instability.
arXiv:2607.17491v1 Announce Type: cross
Abstract: This study provides a quantitative framework for analysis of systemic demand uncertainty and risk propagation cascades across general supply chain networks. By leveraging properties derived from stochastic networks embedded within a Newsvendor paradigm, we model multi-echelon networks under equilibrium and transient operational regimes. We mathematically validate that the systemic volatility behavior commonly referred to as the Bullwhip effect persists entirely as an unavoidable, inherent topological property of coordinated logistics networks, independent of traditional operational noise or information visibility constraints. Extending this paradigm to transient environments, we model inventory drawdown horizons as a multi-dimensional Skorokhod reflection problem. Crucially, we endogenize market-clearing feedback loops by incorporating non-linear price elasticity mechanisms and dynamic trade relation rebalancing, demonstrating how decentralized rational actions co-evolve with physical capacity bottlenecks to accelerate systemic network degradation. Finally, we operationalize the framework through a data-driven numerical experiment mapping global oil trade dynamics, showing how localized chokepoint disruptions, such as a capacity shock in the Strait of Hormuz, trigger non-linear cascading stockouts and systemic reallocation across sovereign buffers over time.
arXiv:2607.17583v1 Announce Type: cross
Abstract: Precision measurement of low-frequency electric field (LFEF) signals with frequency from 30 kHz to 300 kHz is crucial for advancing both fundamental science and practical applications, owing to their unique frequency regime. For conventional electromagnetic antennas, the long wavelength (i.e., several kilometers) of the LFEF leads to a severe size constraint that efficient radiation becomes challenging to achieve when the antenna size is much smaller than the long wavelength of the LFEF signals, which in turn results in a reduction of measurement sensitivity and compromises antenna's performance. By exploiting the high intrinsic sensitivity of cold trapped ions to weak alternating electric signals via Coulomb interaction, we demonstrate a single-ion phonon laser sensor acted by an injection-locked 40Ca+ ion confined in a surface-electrode trap. Combining the beat frequency technique with the injection-locked phonon laser oscillation, we demonstrate a practical and efficient approach for simultaneous extraction of the frequency, phase, and amplitude from a single measurement, without the need for sideband cooling. This approach achieves precision detection for LFEF signals with the sensitivity of 404 uV/(m * Hz1/2) and the detection limit of 61.5 uV/m. Besides, this approach also shows remarkable robustness against noise. Our study helps realizing practical single-atom sensors in the low-frequency regime, opening avenues for applications in subsurface communication, precision metrology, mass spectrometry, and biomedical monitoring.
arXiv:2607.17594v1 Announce Type: cross
Abstract: Commercial fusion energy requires materials that survive intense neutron bombardment whilst extracting extreme heat loads for conversion to electricity. The CuCrZr alloy, the leading heat-sink material for fusion reactors, derives its strength from a fine dispersion of nano-precipitates formed during prime-ageing heat-treatment. Whether this precipitation-hardening strategy can withstand fusion-relevant irradiation remains untested. Here we show, combining in situ transmission electron microscopy under heavy-ion irradiation and He implantation with thermodynamic and transmutation modelling, that the hardening precipitates dissolve under two opposing kinetic regimes: ballistic dissolution dominates at low temperatures, whilst dissolution and re-precipitation dominate at high temperatures. Although the accelerated dose rates inherent to ion irradiation shift the balance between ballistic mixing and thermal back-diffusion relative to reactor conditions, precipitate degradation at both kinetic extremes indicates that the prime-aged microstructure is unlikely to remain unaltered under prolonged neutron exposure. He bubbles and Kr-rich voids nucleate once vacancies become mobile, and transmutation over five service years irreversibly redirects the alloy chemistry towards Ni-Zr intermetallics. These three independent mechanisms converge to challenge the strategy on which CuCrZr performance depends, suggesting that the long-term performance of age-hardenable Cu-based heat-sink alloys in fusion reactors warrants further assessment. Our findings reveal a new materials challenge for fusion reactor design and commercialisation: the need for new Cu-based heat-sink alloys able to retain engineered strength whilst their chemistry is irreversibly rewritten - thermodynamically and ballistically - by the fusion neutron spectrum.
arXiv:2607.17585v1 Announce Type: new
Abstract: Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual tokenizers can discard fine textures and structural details, while separate representation and diffusion training creates a mismatch between reconstruction and generation objectives. These limitations have renewed interest in pixel-space diffusion, which models raw pixels directly, removes the VAE bottleneck, and supports end-to-end optimization. This formulation better matches the demands of high-fidelity generation but introduces challenges in high-dimensional modeling, including noise scheduling, loss weighting, token efficiency, and scalable architecture design. Pixel-space modeling also offers a promising basis for unified multimodal systems: raw pixels, text, and task conditions can be represented in a shared token space and jointly processed by a single Transformer, narrowing the gap between visual understanding and generation. This paper reviews Pixel-Space Diffusion Transformers (pDiTs) from the perspectives of model architecture, continuous generative mechanisms, and unified multimodal modeling. We summarize representative methods, identify key technical challenges, and discuss future directions toward high-fidelity, end-to-end vision foundation models that integrate generation and understanding.
arXiv:2607.16803v1 Announce Type: new
Abstract: Speech Emotion Recognition (SER) is an important component in a wide range of human-centered applications, including healthcare, customer service, and human-omputer interaction. In medical and decision-support settings, there is increasing interest in models that not only achieve accurate emotion recognition but also support transparent predictions and efficient deployment. However, many existing SER approaches rely on complex deep learning architectures that limit interpretability and increase computational cost. This paper presents an explainable and lightweight speech emotion recognition framework based on a compact convolutional neural network architecture. The proposed approach utilizes log-Mel spectrogram representations to capture spectro-temporal speech characteristics and employs attentive statistics pooling to emphasize emotionally salient temporal segments. To improve model transparency, gradient-based class activation mapping (Grad-CAM) is incorporated to visualize the time-frequency regions that influence the model's predictions. Experimental evaluation on the SAVEE emotional speech dataset demonstrates that the proposed framework achieves competitive recognition performance while maintaining a compact architecture with significantly fewer parameters than many existing SER models. The results indicate that efficient convolutional architectures combined with interpretable analysis can provide a practical balance between recognition accuracy, computational efficiency, and model transparency.
arXiv:2607.17099v1 Announce Type: new
Abstract: Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MDE) in both cross-scene generalization and metric-scale prediction, yet these gains have not translated to tiny models. We bridge this gap with DepthART (Depth Anything Rethought for Tiny Models), which is a compact MDE model for on-device deployment across diverse scenes. We first identify two capacity-driven bottlenecks in tiny models: (i) overfitting to dataset-specific distribution bias and (ii) unstable metric adaptation under camera shift, where full fine-tuning easily damages transferable geometry. Accordingly, DepthART combines two simple but effective strategies: a bias-resistant data sampling scheme to reduce distribution bias under the same training budget, and a camera-conditioned fine-tuning protocol that freezes the distilled encoder and adjusts metric scale conditioned on intrinsics while better preserving cross-dataset generalization. Across datasets, DepthART consistently surpasses previous tiny baselines in both zero-shot generalization and metric accuracy (e.g., zero-shot $\delta_1$=0.964 for DepthART-S on NYUD v2), and in some cases approaches heavy models. We further provide a scalable model family, with DepthART-S reaching 347/245 FPS (strict FP32) on an RTX A6000 at $224^2/448^2$, 102 FPS (TF32) on a Orin NX 8GB, and over 15 FPS (FP32) on a Jetson Nano 4GB.
arXiv:2607.16805v1 Announce Type: new
Abstract: High-quality 3D scene assets are critical for embodied applications such as robotic manipulation, navigation, and simulation. Despite their strong object priors, recent single-image 3D generation models such as SAM3D remain insufficient for real-world scenes, where severe occlusions, redundant observations, and cross-view inconsistencies make reliable scene generation challenging. We introduce Scene-SAM3D, a training-free framework that extends SAM3D from single-view object generation to calibrated multi-view scene asset generation. Scene-SAM3D selects a compact set of complementary views, reducing observation redundancy while providing additional evidence for regions occluded in individual views. Based on the selected views, it performs step-efficient latent velocity fusion to integrate multi-view evidence and suppress cross-view conflicts in canonical space. Finally, a lightweight rigid-object Gaussian optimization refines the scene layout within 200 iterations while preserving the generated object geometry. Experiments on Replica and ScanNet++ demonstrate consistent improvements at both instance and scene levels, with our method reducing scene-level CD by 43.8% on Replica and 30.9% on ScanNet++, while cutting flow-model sampling FLOPs and wall-time latency by nearly 20% under the same multi-view setting. Code will be released at https://github.com/xibi777/Scene-SAM3D.
arXiv:2607.17586v1 Announce Type: new
Abstract: Money mule accounts are critical facilitators of financial fraud, yet detecting them at scale remains challenging due to the heterogeneous nature of transactional and behavioural data. We present an end-to-end pipeline for customer-level mule detection comprising three stages: (1) a LightGBM classifier trained on 280 engineered features spanning transaction patterns, account demographics, network topology, and temporal behaviour; (2) a TreeSHAP attribution layer that decomposes each prediction into feature contributions; and (3) a large language model (LLM) module that converts SHAP attributions into analyst-facing natural-language narratives. We evaluate across three open-weight LLM families and assess explanation quality through analyst feedback. In a live production deployment, the system achieves a yield rate of 89%, up from 61% under the incumbent rule-based system, with monthly alert volume expanding from 211 to 302, reflecting broader true-positive coverage rather than increased noise. This corresponds to a 60% incremental adverse detection beyond existing review workflows, substantially outperforming the rule-based approach. Qualitative feedback from analysts indicates that LLM-generated narratives reduce cognitive load during alert triage. We further discuss implications of deploying LLM-augmented explainability in regulated financial environments.
arXiv:2607.17618v1 Announce Type: cross
Abstract: Bell nonlocality is both a defining signature of entanglement and a key quantum information resource. However, visualizing and certifying nonlocal correlations across a spatially multimode photonic field remains challenging due to the rapidly growing measurement cost of spatially resolved projective tests. To address this issue, we build a spatial nonlocality imaging scheme that directly reveals the spatial distribution of quantum nonlocality by integrating a metasurface that performs parallel polarization projections with a quantum-adaptive neural network. Spatially resolved Clauser--Horne--Shimony--Holt (CHSH) tests are realized over a 400-pixel biphoton field using an average of only 1.7 detected coincidence pairs per pixel per basis. This approach yields a nonlocality image that maps the two-dimensional spatial distribution of Bell violations across the optical field and reveals the target-state-dependent spatial evolution of Bell violations. It provides a highly resource-efficient route to large-scale Bell certification and opens new possibilities for exploiting spatially multimode entanglement in quantum imaging, quantum networking, and scalable photonic quantum technologies.
arXiv:2607.17654v1 Announce Type: cross
Abstract: Multidimensional graded response models (MGRMs) are widely used for analyzing ordinal questionnaire data in psychological and educational assessments. A central challenge in applying these models is determining the number of latent dimensions. Conventional approaches usually fit multiple fixed-dimensional models and select among them using post-hoc criteria such as AIC, BIC, or cross-validation, which can be computationally demanding and ignore uncertainty in dimensionality during estimation. We develop an adaptive Bayesian dimension selection framework for probit MGRMs. Building on the cumulative shrinkage process, we assign a cumulative ordered spike-and-slab (COSS) prior to the column-specific variances of the item loading matrix. This prior induces increasing shrinkage across latent dimensions, allowing redundant dimensions to be shrunk toward zero while preserving flexibility for active dimensions. Albert--Chib latent response augmentation is used to handle the ordinal probit likelihood, yielding conditionally Gaussian updates for item loadings and latent traits. These updates are combined with Gibbs updates for threshold and shrinkage parameters in an efficient adaptive sampler.
Simulation studies evaluate the proposed
method in terms of dimension recovery, parameter
estimation accuracy, and computational efficiency, with comparisons to conventional fixed-dimensional estimation and model selection procedures. The results show that the proposed approach accurately recovers the latent structure while avoiding repeated model fitting over multiple candidate dimensions. We further illustrate the method using real psychological assessment data, demonstrating its practical utility for uncovering interpretable latent structures in ordinal item responses.
arXiv:2607.17678v1 Announce Type: cross
Abstract: Monte-Carlo trajectory (quantum-jump) methods are the practical route to simulating noisy quantum circuits once the exact density-matrix method is precluded by its $4^n$ memory cost. Their bottleneck is estimator variance: resolving one expectation value can demand thousands of trajectories. Recent tensor-network work shows that \emph{variance-reduced unravelings} -- projector and analog sampling -- sharply cut this variance, but only on CPU matrix-product-state backends, with no path into production tooling. We implement both unravelings on a \emph{GPU dense-statevector} trajectory engine and validate them against the exact density matrix (ideal-circuit fidelity $1-2.2\times10^{-16}$; $1/\sqrt{N}$ convergence; all unravelings unbiased to trace distance $<0.01$). On a single consumer GPU, projector unraveling reaches a target standard error with $20.8\times$ fewer trajectories than Qiskit-Aer's \texttt{batched\_shots\_gpu} at $n=10$, a factor that holds at $19$--$26\times$ across $n=8$--$20$. A regime map places analog sampling optimal at weak noise and projector at strong noise, crossing near $\gamma t\approx0.35$. We further report a systems finding: Qiskit-Aer applies noise at the \emph{channel} level and reconstructs a canonical Kraus decomposition at apply time, discarding any user-supplied unraveling, so variance-reduced unravelings cannot be delivered through its public API. Because Aer's Born-rule collapse machinery already exists, we specify a minimal change that would unlock the technique in production.
arXiv:2607.17043v1 Announce Type: new
Abstract: Model collapse is a central challenge in learning from synthetic data: as later-generation large language models (LLMs) are trained on an increasing proportion of model-generated data, performance can degrade due to narrowed coverage and accumulated bias. Existing work mainly studies how to bound this degradation. In iterative model evolution, however, the more meaningful objective is to ensure that each successive model improves over its predecessor, which requires diagnosing collapse at a granularity that is actionable for data curation. We study this problem in synthetic data self-improving for instruction tuning. We show that collapse in this setting is not simply uniform performance degradation, but can appear as a polarization of competence, where synthetic training reinforces already strong skills while further degrading weak ones. Motivated by this observation, we propose KITE (Knowledge-boundary Instruction Tuning via Exploration), a two-stage framework that combines failure-guided data generation with boundary-aware uncertainty curation. Experiments across several datasets and multiple open-source LLMs show that KITE yields more stable improvement than strong synthetic-data baselines.
arXiv:2607.17683v1 Announce Type: cross
Abstract: While particle-laden interfaces play a central role in many natural and industrial processes, predicting their mechanical properties remains a major challenge. These systems combine granular characteristics conferred by particle-particle contacts with elastic behavior originating from capillary interactions, making them very sensitive to their history. Using the relaxation of uniaxially compressed particle rafts through a local constriction as a model experiment, we demonstrate the existence of a reproducible and continuous aging process. Aging is observed for both front- and back-compressed rafts and is characterized by a progressive increase in particle mobility and raft deformability. Macroscopic changes are seen, for example, in the extent of relaxation and are correlated with flow modifications observed at the mesoscopic level among which are increased particle fluxes, broader shear zones and enhanced particle rearrangements. While aging can be attributed unambiguously to the constrained passage of the particles through a constriction, its microscopic origin remains hypothetical, the results suggesting that contact lines around the particles may evolve. Beyond providing new insight into the effects of raft history, the proposed constriction flow experiment offers a simple method to control and compare aging in different particulate assemblies.
arXiv:2607.17526v1 Announce Type: new
Abstract: Zero-shot text-guided editing of real-world music recordings requires balancing semantic modification with faithful preservation of the original musical structure. Although recent diffusion transformers trained with rectified flow have achieved remarkable success in text-to-music generation, extending them to edit existing recordings remains challenging because editing requires accurate deterministic inversion, reliable structural preservation, and numerically stable integration throughout the inversion and generation processes.
We present FlowSonic, a zero-shot music editing framework built upon a pretrained diffusion transformer trained with rectified flow. FlowSonic first deterministically inverts a real-world recording into the latent space and preserves its musical structure during editing by reusing cross-attention representations extracted during inversion. To improve the numerical reliability of inversion-based editing, we introduce a high-order ODE solver and systematically investigate how different numerical integration schemes influence trajectory stability, structural preservation, and semantic controllability.
Comprehensive experiments on timbre-transfer and genre-modification tasks demonstrate that FlowSonic consistently outperforms existing music editing methods across semantic alignment, harmonic preservation, structural consistency, and perceptual audio quality. We further provide geometric and empirical analyses showing how the proposed numerical integration strategy improves latent trajectory stability and leads to more reliable music editing.
arXiv:2607.17684v1 Announce Type: cross
Abstract: We study universal monotonicity and Frank--Wolfe stability properties for atomic splittable congestion games. Specifically, we characterize the largest resource cost class for which the associated variational inequality operator is monotone for every game. This characterization is given by a curvature inequality involving the first two derivatives of the allowable cost functions and the number of players. Our framework yields exact characterizations for universal monotonicity, strict monotonicity, and strong monotonicity; the strict and strong variants require corresponding stricter curvature conditions. We then draw a perhaps surprising connection to learning dynamics in atomic splittable congestion games. We show that the very same curvature condition also characterizes universal \emph{local and global stability} of the Euclidean-regularized Frank--Wolfe dynamics on arbitrary convex strategy spaces, provided the cost class is closed under positive affine transformations. Finally, we study games on simplices and show that an interior equilibrium of the regularized Frank--Wolfe dynamic is locally exponentially stable, even without the curvature condition.
arXiv:2607.16669v1 Announce Type: new
Abstract: OpenLanguageModel (OLM) is an open-source PyTorch library for building and pretraining small language models while keeping their machinery visible. In OLM, model code reads like the architecture: components are ordinary modules, while Block, Residual, Repeat, and Parallel describe how they are wired. The resulting model can move unchanged from a teaching notebook to a complete pretraining run or a research ablation. OLM connects this readable model layer to tokenizers, local and streaming datasets, optimization, mixed precision, callbacks, checkpoints, and hardware-aware CPU, single-GPU, and single-node multi-GPU execution. We demonstrate the full path by tracing GPT-2 from diagram to code, launching a FineWeb-Edu training script, replacing one attention component, and letting AutoTrainer configure the available machine. The package includes 27 presets across nine familiar model families and documentation that progresses from LM fundamentals to architecture research. Validation shows close agreement with independent reference implementations, 90.6% four-GPU weak-scaling efficiency for a 348M-parameter workload, compact architecture edits, and positive early usability results. OLM is MIT-licensed and available through PyPI, GitHub, and its documentation site.