arXiv:2607.09392v1 Announce Type: new
Abstract: We consider the inverse problem for the quantitative reconstruction of physical properties in the context of passive imaging, where ambient wavefields are used to infer a medium. The data are modeled as a superposition of waves generated by stochastic sources. In this work, we focus on time-harmonic acoustic wave propagation and assume that the stochastic sources exciting the medium are zero-mean and spatially uncorrelated. Under these assumptions, the expected value of the cross-correlation between signals recorded at two locations can be related to the deterministic Green's function and the covariance of the source terms. We follow a first-order formulation of the wave equation, which enables the treatment of correlations between different types of wavefields. A numerical framework is developed for the resulting nonlinear inverse problem. The quantitative reconstruction is carried out using an iterative minimization scheme, in which the gradient of the misfit functional is computed via the adjoint-state method. Numerical experiments in two and three dimensions are performed using synthetic data, and inversions based on the expected value of cross-correlations are compared with those relying on direct wavefield measurements from active-source acquisitions.
Science Journals
arXiv:2511.09524v2 Announce Type: replace
Abstract: The concept of a security index quantifies the minimum number of components that must be compromised to carry out a stealth attack. This metric enables system operators to assess the security risk of each component and implement countermeasures accordingly. In this paper, we introduce a data-driven security index that can be computed solely from input/output data when the system model is unknown. We show a sufficient condition under which the data-driven security index coincides with the model-based security index, which implies that the exact risk level of each component can be identified solely from data. We also provide an algorithm for computing the data-driven security index.
arXiv:2607.08853v1 Announce Type: new
Abstract: Mobility systems of people and goods are inherently multi-scale, spanning levels of organization from individual cities to regions and nations. Understanding whether mobility networks exhibit similar patterns across these scales is important. Such similarity would point to common organizing principles, enabling insights gained at one scale to inform planning and management at others. Despite growing efforts to analyze mobility at multiple scales, such cross-scale similarity remains poorly understood, and renormalization provides a natural framework for addressing this question. Here, we propose a Neighbor-Limited Box Covering method to renormalize undirected weighted mobility networks. This method iteratively selects box centers in descending order of node strength, merges each center with a fixed number of its highest-weight neighbors to form a renormalized node, and aggregates edge weights between renormalized nodes to generate the network at the next scale. We apply this technique to uncover multi-scale structures of real-world inter-city human mobility and freight trip networks in China and find that the topological structures, weighted structural features, and dynamic processes all exhibit self-similarity across these multi-scale mobility networks. Moreover, we find that the constituent nodes in most renormalized nodes show a strong spatial cohesion, and the boundaries of them closely follow existing political and socio-economic borders, even though the method does not explicitly incorporate any spatial information. Our study not only reveals the consistency of multi-scale inter-city mobility patterns, but also provides important insights into their spatial organization. Furthermore, our method is applicable to mobility networks of different sizes and has potential as a powerful tool for the multi-scale analysis of various other real-world complex systems.
arXiv:2607.09399v1 Announce Type: new
Abstract: We introduce a novel method for both partial and full optimization of the connections in deep differentiable logic gate networks (LGNs) and lookup table networks (LUTNs). Our training method utilizes a probability distribution over a set of connections per gate/lookup table (LUT) input pin, selecting the connection with highest merit, all whilst the optimal gate types or LUT-entries are learned in parallel. We show that the connection-optimized LGNs outperform standard fixed-connection LGNs on the Yin-Yang, MNIST Handwritten Digits and Fashion-MNIST benchmarks, while requiring only a fraction of the number of logic gates. We achieve 98.92% on the MNIST dataset with two layers of 8000 gates. With only one layer of 8000 gates, we obtain 98.45%, showing that our method requires almost 50 times fewer gates compared to fixed-connection LGNs. Training stability up to ten layers has been ensured by employing a high learning rate, straight-through estimators and trimming constant-output gate types. Additionally, we present a LUT neuron description that enables stable training with backpropagation, tested up to 6-layer deep networks. The model requires four times fewer trainable parameters and still achieves a higher accuracy compared to the fixed-connection LGN training algorithm. Our connection-training algorithm also works well for the LUTNs, achieving an accuracy of 98.88% for two layers of 2000 6-input LUTs.
arXiv:2607.08859v1 Announce Type: new
Abstract: Visibility in media is pivotal for identity development and for broadening societal views of gender and sexuality. Queer representation has increased in recent years, yet damaging stereotypes and tropes persist. Here, we focus on queer portrayal and its perception by audiences in fictional stories (television, film, and literature) by studying characters by their quantified archetypes which are operationalizations of common conceptions such as Hero, Diva, and Outcast. We use the archetypometrics and Fandom's LGBTQIA+ datasets to study samples of fictional characters along the trait differential spanning straight to queer. We find, quantify, and explain a seeming paradox. The characters with the highest queer score present positive primary archetypes and are typically Heroes rather than Fools, Angels rather than Demons, and Adventurers rather than Traditionalists. But evaluation across many stories for the straight-queer trait itself reveals a strong collective-writing bias towards Fool (away from Hero) and no meaningful loading for the other two dimensions. Our analysis offers a population-scale view of the complexities of queer portrayal, while also pointing to risks in blindly training on many-authored story corpora.
arXiv:2604.01480v2 Announce Type: replace
Abstract: Metasurface inverse design can realize complex optical functionality, but turning a target optical response into executable optimization code still requires substantial expertise in computational electromagnetics and solver-specific software engineering. We present a self-evolving agentic framework that lowers this barrier by coupling a coding agent, explicit human-readable skill files, and a deterministic physics-based evaluator. Rather than updating model weights, it revises the skill files from solver-grounded feedback, while the base model and differentiable solver, which provides the physics simulation and gradients, stay fixed. On a multi-type benchmark, skill evolution raises same-type task success from 38\% to 74\%, the fraction of physical criteria met from 0.51 to 0.87, and reduces average attempts from 4.10 to 2.30. On two new-type families, success holds near ceiling on one (0.92 to 0.90) and rises from 0.20 to 0.90 on the other. Skill evolution offers a practical path toward autonomous and accessible inverse-design workflows.
arXiv:2607.08863v1 Announce Type: new
Abstract: We present Clean2FX, a study and demo of label-conditioned clean-to-effect transformation for electric guitar audio. Given a clean guitar input and a target effect label, the task is to synthesize the corresponding effected signal while preserving the musical content. Training and evaluation pairs are constructed from EGFxSet real, single tone recordings by assembling matched clean/effected chords, melodies, and mixed timelines. This allows for controlled comparison across effects. We evaluate four neural approaches under a common spectrogram-based transformation setting: two variational autoencoders and two U-Net models that differ in whether they operate on linear or log-magnitude representations. Performance is measured using linear-magnitude spectrogram MSE and Fr\'echet Audio Distance. The U-Net models outperform the variational autoencoder variants. Per-effect results show that distortion effects are most readily improved, whereas delay and reverb effects exhibit weaker FAD gains despite substantial spectral-error reductions. A conditioning-sensitivity diagnostic provides evidence that the best model responds to target labels rather than collapsing to a single transformation. Our demo website compares two models applied on real-world guitar performances outside training and validation data, providing audio and spectrogram examples of the practical clean-to-effect behavior.
arXiv:2601.21859v2 Announce Type: replace
Abstract: The fundamental trade-off between privacy and utility remains an active area of research. Our contribution is motivated by two observations. First, privacy mechanisms developed for one-time data release cannot straightforwardly be extended to sequential releases. Second, practical databases are likely to be useful to multiple distinct parties. Furthermore, we can not rule out the possibility of data sharing between parties. With utility in mind, we formulate a new privacy-utility trade-off problem to adaptively tackle sequential data requests made by different, potentially colluding entities. We consider both expected distortion and mutual information as measures to quantify utility, and use mutual information to measure privacy. We assume an attack model whereby illicit data sharing, which we call collusion, can occur between data receivers. We develop an adaptive algorithm for data releases that makes use of a Blahut-Arimoto-style algorithm. We show that the resulting data releases are optimal when expected distortion quantifies utility, and locally optimal when mutual information quantifies utility. Numerical experiments on real data demonstrate that the proposed adaptive algorithm can exploit previously released information to reduce cumulative leakage under collusion without sacrificing much, if any utility. Finally, we discuss how our findings may extend to applications in machine learning.
arXiv:2607.09562v1 Announce Type: new
Abstract: Medical Vision-Language Models (VLMs) exhibit strong zero-shot performance, yet their effectiveness still declines on out-of-distribution (OOD) data due to domain shifts and class bias inherited from large-scale pretraining. Existing few-shot adaptation methods typically introduce additional trainable components, which can be unstable in extremely low-data regimes (e.g., 1-shot), and lack robustness on different medical data. We present TCLA, a purely training-free few-shot adaptation method for Medical VLMs, which is fast and model-agnostic. TCLA corrects inference logits based on a small set of support samples, boosting pretrained VLMs performance by improving inter-class deconfusion and reducing domain shift. Extensive experiments on nine datasets across multiple medical imaging modalities including X-ray, Ultrasound, MRI, CT, Histopathology, demonstrate that TCLA consistently improves OOD performance of Medical VLMs and, in most of cases, outperforms existing training-based adaptation methods.
arXiv:2607.09020v1 Announce Type: cross
Abstract: Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both tasks. We leverage S3M-based Phonological Activation Mapping (SPAM), which maps each S3M representation frame to a vector of phonological feature activations, such as voicing and nasality. On top of SPAM, we introduce two simple but effective lightweight, gradient-descent-free prediction heads: a recognition head and a segmentation head. Our method requires less than a minute of phonetic transcriptions, and generalizes to unseen phones during training. Across a diverse range of datasets, our approach attains strong segmentation and recognition performance.
arXiv:2604.22294v2 Announce Type: replace
Abstract: Systematic reviews -- which requires comprehensive evidence collection and synthesis from large document corpora in response to targeted research questions -- are foundational in finance, social sciences, and other technical fields. Manual construction of evidence tables is labor-intensive, and recent LLM-based assistants relying on embedding or keyword based search often fail to meet the coverage standards of systematic reviews. We introduce SLIDERS, a novel LLM-based methodology for systematic reviews, by automatically assembling evidence tables tailored to research questions. In addition to extracting structured data from documents, SLIDERS can extract full-text excerpts that serve as direct evidence or as provenance for structured data. Core to SLIDERS is an automated evidence reconciliation agent that writes code to analyze and reconcile extracted evidence, bringing together information fragmented across documents, resolving inconsistencies across excerpts, and synthesizing overlapping findings into a coherent evidence table. In addition, SLIDERS allows users to ask follow-up questions in natural language to further explore the assembled evidence. We evaluate SLIDERS on three systematic-review-style tasks over large document collections. SLIDERS outperforms the best-performing baseline across benchmarks, remains near 90% accuracy across 6M-11M-token corpora. On two new follow-up analysis benchmarks SLIDERS can answer 77.9% and 58.3% followup questions accurately
arXiv:2607.09167v1 Announce Type: new
Abstract: Schedule-Free methods have attracted growing interest for alleviating the burden of designing and tuning a learning rate scheduler, while matching and sometimes even outperforming optimizers with tuned schedulers. Despite their strong empirical results, their convergence theory in nonconvex optimization, where modern machine learning objectives typically arise, has remained largely unexplored. In this paper, we provide worst-case analyses of Schedule-Free gradient descent and Schedule-Free stochastic gradient descent, in their standard form and without auxiliary modifications or restrictive conditions, for smooth but possibly nonconvex objectives. Based on a Lyapunov analysis derived from the continuous-time limiting ordinary differential equation associated with these methods, we show that Schedule-Free gradient descent and Schedule-Free stochastic gradient descent achieve the optimal worst-case convergence rates attainable among first-order methods. We further formulate Schedule-Free gradient descent as a nonautonomous dynamical system and prove strict-saddle avoidance under an arbitrarily small one-time perturbation. These theoretical results provide a better understanding of the strong performance that Schedule-Free methods demonstrate.
arXiv:2601.06763v2 Announce Type: replace-cross
Abstract: The motion of atoms in programmable optical tweezer arrays offers many new opportunities for neutral atom quantum science. These include inter- and intra-site atom motion for resource-efficient implementations of fermionic and bosonic modes, respectively, as well as tweezer transport for efficient compilation of arbitrary circuits. However, the exploitation of atomic motion for all three purposes and others is limited by the inertia of the atoms. We present a comprehensive architectural blueprint for the use of fermionic metastable helium-3 ($^3$He$^*$) atoms -- the lightest trappable atomic species -- in programmable optical tweezer arrays. This includes a concrete analysis of atomic structure considerations as well as Rydberg-mediated interactions. We show that inter-tweezer hopping of $^3$He$^*$ atoms can be $\gtrsim3\times$ faster than previous demonstrations with lithium-6. We also demonstrate a new toolbox for encoding and manipulating qubits directly in the tweezer trap potential, uniquely enabled by the light mass of $^3$He$^*$. Finally, we provide several examples of new opportunities for fermionic quantum simulation and computation that leverage the transport and inter-tweezer hopping of $^3$He$^*$ atom arrays. These tools present new methods to improve the resource efficiency of neutral atom quantum science that may also enable quantum simulations of lattice gauge theories and quantum chemistry outside the Born-Oppenheimer approximation
arXiv:2607.09169v1 Announce Type: new
Abstract: Egocentric 3D human pose estimation from head-mounted stereo cameras is challenging due to fisheye distortion, severe self-occlusion, and frequent truncation of body joints outside the camera field of view. Recent stereo egocentric methods have improved performance through heatmap lifting, stereo correspondence, and transformer-based refinement, but they often rely heavily on frame-local evidence or use temporal information only as auxiliary pose-level context. This limits robustness when current-frame stereo cues are weak, occluded, or ambiguous. We propose TSR-Ego, a temporally guided stereo framework that couples short-term motion evidence with projection-guided feature sampling. The model first enriches dense stereo feature maps using a causal depthwise-separable temporal convolution, allowing past visual evidence to influence the feature space before deformable cross-attention. A single-stage causal stereo decoder then refines learned 3D joint queries through temporal self-attention, joint self-attention, and fisheye deformable stereo cross-attention, using the evolving pose estimate to generate 2D sampling references. Unlike methods that apply temporal reasoning mainly after pose prediction, TSR-Ego uses motion context to shape both the sampled stereo features and the joint representations while preserving online inference without future frames. Experiments on UnrealEgo2 and UnrealEgo-RW show state-of-the-art performance, with especially strong gains on real-world sequences.
arXiv:2607.09097v1 Announce Type: cross
Abstract: We study stochastic fixed-point equations $\mathbf{T}(\mathbf{x}) = \mathbf{x}$ over normed spaces $(\mathcal{E}, \|\cdot\|)$, where the operator $\mathbf{T}$ is nonexpansive or contractive and is accessed only through unbiased stochastic evaluations with bounded second central moment. Given $\epsilon > 0, \delta \in (0, 1)$, the goal is to output $\mathbf{x} \in \mathcal{E}$ such that $\|\mathbf{T}(\mathbf{x}) - \mathbf{x}\| \leq \epsilon$ with probability at least $1-\delta$. We introduce VR-GHAL, a variance-reduced gradual Halpern method for quadratically smoothable Banach spaces. The key algorithmic ingredient is a recursive stochastic estimator based on clipped differences of oracle evaluations: instead of clipping $\tau(\mathbf{x}; \xi)$ itself, we clip stochastic differences at the Lipschitz scale $\gamma\|\mathbf{x} - \mathbf{y}\|$. This makes the estimator pathwise Lipschitz along the algorithmic trajectory while permitting martingale concentration under finite second moments in the native norm. Our main theorem gives an anytime high-probability residual bound: on a single event of probability at least $1 - \delta$, the residual decreases nearly geometrically across epochs, up to lower-order logarithmic factors. Under only bounded variance, displaying only the dependence on the target error $\epsilon$ and Lipschitz constant $\gamma \in (0, 1]$ of $\mathbf{T}$, the resulting oracle complexity is $\min\{\epsilon^{-5}, (1-\gamma)^{-3}\epsilon^{-2}\}$. Under a Lipschitz-in-expectation oracle, the dependence improves to the corresponding $\epsilon^{-3}$ nonexpansive rate (i.e., for $\gamma = 1$), and under samplewise nonexpansiveness to $\epsilon^{-2}$.
arXiv:2607.09139v1 Announce Type: cross
Abstract: We study the optimal transport of optimally controlled agents from a compactly supported absolutely continuous source to a discrete target measure. The ground cost for the transport is induced by the optimal cost of the agents' motion. When this ground cost satisfies the twist condition, the optimal transport map is given almost everywhere in terms of a Laguerre tessellation of the state space. We refer to this control-theoretic generalization of Laguerre tessellation as Control Laguerre Tessellation (CLT), and illustrate it for two ground costs induced by linear controlled agents with minimum energy and minimum time objectives.
arXiv:2511.17813v3 Announce Type: replace
Abstract: LLM-based simulations can enable controlled studies of civic deliberation, but current systems lack speaker-attributed data and methods for evaluating long-form institutional behavior. ASR transcripts typically use anonymous labels such as $Speaker\_1$, preventing models from learning stable participant behavior across meetings. We present a reproducible pipeline that converts public Zoom recordings into speaker-attributed transcripts enriched with persona profiles, topics, and pragmatic "action tags" such as $[propose\_motion]$. Using this pipeline, we release three public datasets of government deliberation (Appellate Court hearings, School Board meetings, and Municipal Council sessions) and fine-tune LLM personas on this action-aware data. We evaluate simulations along four dimensions: persona fidelity, persona consistency, institutional fidelity, and behavioral coherence. Action-aware fine-tuning cuts perplexity by 67%, doubles classifier-based persona fidelity, increases vote attempts by up to $3.6\times$, and improves deliberative responsiveness by up to 70%. Human evaluations show that simulated excerpts are often hard to distinguish from real deliberations, indicating a practical foundation for data-grounded civic simulation studies.
arXiv:2607.08870v1 Announce Type: new
Abstract: The well-known Shallow Water Equations (SWE) are used for modeling incompressible free-surface flows whenever the shallowness allows for a vertical-averaging; i.e., vertical effects are negligible in comparison to horizontal ones. But vertical averaging comes with the price of losing information along the vertical axis. Moment models for shallow flow contain information on the vertical velocity and pressure profile despite being dimensionally reduced. A class of these models incorporating a non-hydrostatic pressure have been introduced before as Dispersive Shallow Moment Models (DSM). However, no method for solving the non-stationary equations has been presented yet, mainly because it was unclear how to compute the pressure equation in the form of the divergence-free constraint. We rewrite the pressure equations of the DSM models in the form of a Poisson-like problem to enable their solution with a projection-type splitting scheme. For the linear equations, we present the calculations for the generalized model and discuss the non-linear case. We state the first two linear models and the corresponding nonlinear counterparts. Finally, we introduce a hybrid Finite-Volume Finite-Difference method and discuss the non-stationary numerical results for an experiment with periodic boundary and uneven bottom topography.
arXiv:2607.09403v1 Announce Type: new
Abstract: Worldbuilding, the construction of coherent fictional worlds, is a foundational task in game design and literary creation. Large Language Models (LLMs) offer new possibilities for automated content generation, but their application to worldbuilding faces three challenges: context explosion that grows linearly with the building process, the tension between creative diversity and content consistency, and the absence of automated quality assurance. This paper presents AutoWorldBuilder, a multi-agent collaborative system that addresses these challenges through five integrated components: a structured concept network with conflict detection; a DAG-based hybrid batch scheduler that groups tasks by semantic locality; a four-layer context compression mechanism achieving approximately 90% token reduction; an iterative review system with specialized Auditor agents that improves proposal pass rates from 42% to over 85%; and a skill-driven agent architecture supporting zero-code extension with differentiated temperature configuration. Two experiments across 20 diverse worldbuilding tasks, using GPT-OSS 120B and DeepSeek v3.2 as LLM backends, demonstrate a 95.0% success rate. The system generated 56-103 self-consistent concepts per world in 18-31 minutes with zero-conflict delivery. The architectural patterns validated here, including layer-as-budget compression, semantic-locality scheduling, and separation of generation and review, transfer to the broader class of knowledge-intensive, multi-agent LLM applications.
arXiv:2508.14817v2 Announce Type: replace
Abstract: Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs). Methods: We defined three EHR-based tasks that are replicable across health systems and vary in reasoning complexity: 1) extracting imaging procedures (modality, date, and anatomic site), 2) generating timelines of therapeutic antibiotic use, and 3) identifying the key diagnoses for a hospitalization. Using real inpatient clinical notes from a US academic health system, we evaluated three large language models (GPT-5.4-mini, Mistral Medium 3, DeepSeek V3.1) with varying amounts of provided context, comparing targeted retrieval to using the most recent clinical notes. Results: For Imaging Procedures, RAG strongly outperformed recent-note inputs and exceeded long-context performance (by 0.17-9.83 F1 across all models) using fewer than 8K tokens. Similar benefits were observed for Antibiotic Timelines, where <8K of retrieved tokens matched long-context recent-notes performance (between -3.26 to +3.24 Jaccard). Error analysis revealed that missing information in the clinical notes--often due to inter-hospital transfers--limited performance to some extent. However, performance on the Diagnosis Generation task remains largely static across methods and models. Discussion: RAG demonstrated strong token efficiency across tasks, with the clearest and most consistent gains observed for imaging extraction and antibiotic timeline reconstruction. Diagnosis generation proved the most challenging task, suggesting ceiling effects imposed by documentation variability and evaluation constraints. Conclusion: Our results suggest that RAG remains a competitive and efficient approach for clinical tasks over large amounts of EHR, even as newer models become capable of handling increasingly longer amounts of text.
arXiv:2512.03238v2 Announce Type: replace
Abstract: High quality data is needed to unlock the full potential of AI for end users. However finding new sources of such data is getting harder: most publicly-available human generated data will soon have been used. Additionally, publicly available data often is not representative of users of a particular system -- for example, a research speech dataset of contractors interacting with an AI assistant will likely be more homogeneous, well articulated and self-censored than real world commands that end users will issue. Therefore unlocking high-quality data grounded in real user interactions is of vital interest. However, the direct use of user data comes with significant privacy risks. Differential Privacy (DP) is a well established framework for reasoning about and limiting information leakage, and is a gold standard for protecting user privacy. The focus of this work, \emph{Differentially Private Synthetic data}, refers to synthetic data that preserves the overall trends of source data,, while providing strong privacy guarantees to individuals that contributed to the source dataset. DP synthetic data can unlock the value of datasets that have previously been inaccessible due to privacy concerns and can replace the use of sensitive datasets that previously have only had rudimentary protections like ad-hoc rule-based anonymization.
In this paper we explore the full suite of techniques surrounding DP synthetic data, the types of privacy protections they offer and the state-of-the-art for various modalities (image, tabular, text and decentralized). We outline all the components needed in a system that generates DP synthetic data, from sensitive data handling and preparation, to tracking the use and empirical privacy testing. We hope that work will result in increased adoption of DP synthetic data, spur additional research and increase trust in DP synthetic data approaches.
arXiv:2508.15806v2 Announce Type: replace
Abstract: The growing sequence length of large language models poses significant challenges for key-value (KV) caches. Existing state-of-the-art cache eviction methods primarily analyze the inference behavior of attention heads in successful retrieval-reasoning cases, often overlooking diverse behaviors in failure cases, such as bias and distraction. This oversight limits the potential to leverage heterogeneous head behaviors for improved eviction performance. Inspired by the confusion matrix, we introduce an Attention Behavior Matrix to comprehensively analyze attention head behaviors in both success and failure scenarios. By maximizing the signal-to-noise ratio -- strengthening valid reasoning pathways in success cases while inhibiting noise from bias and distraction in failure cases -- we propose REtrieval-reAsoning and Logic-constructed (REAL) KV cache eviction, the first method to leverage multi-behavior analysis. Comprehensive evaluations show that REAL achieves remarkable performance across various models and benchmarks; notably, on LongBench v2, it achieves comparable accuracy to the strongest baseline, HeadKV-R2, while requiring 32x less space (Figure 1). By offering a novel perspective on behavior analysis, we pave the way for a shift from success-only to comprehensive, failure-aware methods in long-context modeling. Our code is available at https://github.com/yonseicasl/REAL.
arXiv:2607.08878v1 Announce Type: new
Abstract: Experiments on the regeneration of long streaks in flows in which they had originally been damped show that their initial growth is due to the interaction of the mean shear with long cross-flow velocities (rollers) that remain even when the streaks are damped. More surprisingly, turbulence also persists in simulations in which only the long rollers are damped while long streaks remain, and these flows are also able to recover when the damping is removed. Finally, simulations are presented in which both streaks and rollers longer than $\lambda_x^+ \approx 600$ are damped. They survive and regenerate, and an interpretation in terms of their energy balance is provided. In contraposition to the classical minimal channels, which include infinitely long structures, these new flows do not contain features longer than the damping wavelength, and support a model in which wall turbulence only depends on processes for which the geometric aspect ratio is of order unity.
arXiv:2603.00620v2 Announce Type: replace
Abstract: Multilingual NLP research increasingly involves hundreds or thousands of languages across different datasets. Managing, discovering, and reporting language metadata becomes a common hurdle at these scales. We present QQ, a metadata toolkit and browser explorer. QQ compiles language metadata sources into a graph of language varieties, scripts, regions, identifiers, names, and relations, and exposes it through a Python API, a command-line interface, and a browser-based explorer. Users can normalize identifiers, retrieve metadata, traverse relations, and discover which external resources contain a language. We demonstrate QQ on three workflows: an audit of the HuggingFace Hub, linking resources that use different identifier systems, and generating reproducible language-reporting tables. QQ supports FAIR-oriented metadata practices through versioning, open formats, and reusable interfaces.
arXiv:2607.09184v1 Announce Type: cross
Abstract: Zak-Orthogonal Time Frequency Space (OTFS) modulation is known to be robust to Doppler spread in high mobility scenarios when compared to Orthogonal Frequency Division Multiplexing (OFDM). This is due to the fact that the channel response to a Zak-OTFS carrier within a frame can be accurately estimated from the channel response to another carrier within the same frame. However, an important open problem and question is whether inter-frame channel prediction is possible with Zak-OTFS, i.e., is it possible to accurately predict the channel response to a Zak-OTFS carrier in a frame based on knowledge of the channel response to some Zak-OTFS carrier in \emph{another} frame (i.e., not the same frame).
In this paper we show that indeed inter-frame channel prediction is possible. We show that the effective DD domain channel filter coefficients vary in a deterministic manner as we move from current to future frames in time and frequency. We also show that the subspace spanned by channel filter coefficients of consecutive frames in time/frequency is invariant to discrete shifts in time and frequency. We exploit the deterministic variation and subspace invariance to propose a novel deterministic ESPIRIT-type method which uses the effective DD domain channel filter taps/coefficients estimated in training frames (i.e., current/past frames in time and frequency having both pilot and data carriers) to predict the effective DD domain channel filter for frames which are several tens of frames in future and several tens of frames away in frequency.