Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

It is not enough to give your moderation rules to ChatGPT: Policy-as-Prompt Moderation and Its Potential Impacts on Community Governance
arXiv:2607.12149v2 Announce Type: replace Abstract: Content moderation practices and governance paradigms are changing rapidly, as fewer human moderators are deployed as `experts' by social media companies in a centralized manner. Instead, the companies are focusing more on community approaches, relying on volunteers to provide accurate information and make correct decisions. In decentralized moderation, communities have always relied on volunteers, updated community guidelines, and internal discussions thereof. For both content moderation paradigms, Artificial Intelligence (AI) seems like it could help ease moderation burdens of time, mental health, and accuracy. One possible way to operationalize AI in content moderation is a `policy-as-prompt'' approach, where the policy is formulated as a natural-language prompt and then passed to a large language model (LLM). This model then aids in moderation tasks. In this paper, we briefly lay out the technical and governance properties of this approach, and argue that its limitations lead to specific risks and harms that have to be addressed. Towards alleviating them, we lay out multiple considerations towards more effective prompt governance, but ultimately find that writing prompts alone is not appropriate for ensuring meaningful community governance.
VTM-Nav: Hierarchical Visual-Topological Memory for Cross-Episode Object-Goal Navigation
arXiv:2607.14514v1 Announce Type: new Abstract: Object-goal navigation requires an embodied agent to locate and reach an instance of a specified object category in an indoor environment. Recent training-free approaches leverage vision-language models (VLMs) for open-vocabulary semantic reasoning, but are typically evaluated under an episodic protocol that resets all scene-specific state after each episode. We introduce Cross-Episode Object-Goal Navigation, in which an agent repeatedly operates in the same scene, retains only self-acquired experience, and keeps its model parameters fixed. To support experience reuse, we present \method, a training-free VLM navigation framework with a persistent hierarchical Visual-Topological Memory (VTM). The VTM organizes scene knowledge at room and object levels and retrieves relevant experience through coarse-to-fine matching, providing memory as soft guidance only when it agrees with current observations. A conservative execution guard further mitigates oscillations, blocked motions, and premature stopping. Under a controlled same-scene protocol, we evaluate \method{} on three benchmarks, HM3D v0.1, HM3D v0.2, and MP3D, and compare it with a strengthened WMNav baseline augmented with cross-episode textual memory, while keeping the VLM backbone and action pipeline identical. \method{} achieves the best performance across all three benchmarks, demonstrating the effectiveness and robustness of structured visual-topological experience reuse across datasets.
Photovoltaics beyond the panel: system-level net energy under grid and firming constraints
arXiv:2607.14784v1 Announce Type: new Abstract: Photovoltaics can exhibit favourable installation-level energy return on investment (EROI), yet electricity must also be delivered when solar and wind generation are insufficient. This study evaluates how the energy balance changes when the boundary is expanded to reliable electricity supply. A fleet-equivalent benchmark model includes component turnover, grids, storage, renewable overbuild, dispatchable backup, and fuel supply. It is not a simulation, optimization, or real-system prediction. Photovoltaics provide the accounting reference in an illustrative Central European mix of 80% wind and 20% photovoltaics, while the firming penalties arise more generally from weather-dependent generation. Expanding the boundary is decisive. The unfirmed renewable fleet retains an EROI of 10.7-16.0, but does not provide continuous electricity. Battery load-balancing references decline from 7.1-12.5 for 1h to 0.7-2.3 for 24h. Seasonal hydrogen yields 3.4-7.7; adding a 1h battery reduces the combined range to 2.9-6.8. Pipeline-gas and LNG backup yield only 1.2-2.6 and 1.0-2.2. The lifecycle-carbon advantage narrows accordingly. Under present industrial supply chains, GWP$_{100}$ intensities rise from 38-75 gCO$_2$eq/kWh without firming to 80-289 gCO$_2$eq/kWh for hydrogen and battery-hydrogen firming, and 121-327 gCO$_2$eq/kWh for gas backup. These favourable benchmarks set orientation, site, curtailment, and delivery factors to unity; real deployment lowers EROI and raises emissions. The results identify firming as a critical energetic constraint: reliable decarbonization requires firm low-carbon generation with sufficient net-energy surplus to sustain a resilient industrial society.
Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion
arXiv:2607.14371v1 Announce Type: new Abstract: Large Language Models (LLMs) have revolutionized AI services, but a critical tension emerges: while personalization improves model performance, it consumes scarce computational resources that users must share. When should a user invest in expensive Supervised Fine-Tuning (SFT) versus lightweight In-Context Learning (ICL)? How does congestion from other users' personalization choices reshape these incentives? And what strategies should platforms adopt when offering multiple personalization algorithms? We develop a tractable framework for LLM serving that captures the statistical-economic trade-offs users face. Our analysis yields several surprising insights. First, we show that ICL and SFT dominate in different regimes, determined by an interplay between pretraining coverage and data signal-to-noise ratios, but congestion can flip these rankings. Second, equilibrium resource consumption exhibits pronounced non-monotonicity: improving pretraining precision reduces the congestion, while broader pretraining coverage and harder tasks sometimes increase it. Third, we prove that offering both personalization methods never hurts the platform's maximal profits, despite potentially increasing computational load. Experiments with GPT-2 on linear regression tasks validate our theoretical predictions about algorithm performance. Complementing these results, our review of documentation from 21 major AI platforms shows that the share offering both SFT and ICL increased from 9.5% in 2021 to 71.4% in 2025, consistent with our platform-design implications.
CASP: Learning-Augmented Offline Approximation with Verifiable Certificates and Bounded-Loss PAC Guarantees
arXiv:2607.14545v1 Announce Type: new Abstract: Machine-learned predictions can speed up offline NP-hard optimization, but asking a predictor what to do amounts to asking it to solve the problem, and committing an unchecked prediction forfeits every worst-case guarantee. CASP (Certificate-Augmented Solution Pruning) instead asks which parts of the search space may be ignored, and accepts each answer only after a sound polynomial-time verifier has checked it, so correctness never depends on prediction quality. We develop the learning theory of this design. The verifier makes the induced loss class uniformly bounded, so certificate parameters are learnable from $\tilde O(\varepsilon^{-2}\log K)$ samples ($K$ the maximum instance size), whereas the unverified commitment class admits no distribution-free rate and, under cost spread $R$, none below $\Omega(R/\varepsilon^2)$. Filtering noisy predictions by verifiable confidence dominates the standard min-combiner, with a margin we compute in closed form, and the prediction stays useful even given the LP, because it breaks ties on degenerate optimal faces, where every symmetric LP policy, meaning one whose commitments depend on the instance only through the verifiable confidence values, provably stalls. Experiments on five problems test the theory's quantitative predictions. With trained predictors, unverified pruning loses up to $26%$ of the optimum under distribution shift, while the verified deployment of the same predictions loses nothing.
Inferring Non-Normal Amplification Geometry from Multivariate Time Series
arXiv:2607.14786v1 Announce Type: new Abstract: Across hydrodynamics, ecology, neuroscience, network dynamics, non-Hermitian physics, and socio-economic systems, asymptotically stable dynamics can exhibit large transient amplifications that are invisible to eigenvalue-based analyses. The mechanism is geometric rather than spectral: perturbations entering along one direction may be expressed transiently along another, allowing asymptotic decay to coexist with strong transient or noise-driven amplification. We introduce non-normal directional response inference, a data-driven method for detecting this geometry from multivariate time series when the governing operator is unknown. A local linear operator is estimated from sliding windows and projected onto the dominant two-dimensional input-response subspace. The reduced dynamics are summarized by the eigenvalue splitting $\Delta$, eigenvector non-orthogonality $K$, and the scale-free ratio $R=K/K_c(\Delta)$, where $K_c(\Delta)$ is the two-dimensional threshold for transient amplification. Controlled benchmarks show that the reduced geometry, particularly $R$, can be recovered from finite data even when the full high-dimensional operator is poorly estimated. Tests across sample size, dimension, training horizon, spectral structure, and non-stationarity confirm that the relevant response geometry requires far fewer observations than full-matrix recovery. Applied in moving windows to electrohysterogram, seizure EEG, freezing-of-gait, and unstable push-up inertial recordings, the method reveals systematic changes around known physiological or behavioral episodes through shifts in $R$, changes in $\Delta$, or stronger projection of fluctuations onto the inferred response direction. It thus exposes interpretable changes in local response geometry without framing the problem as supervised event detection.
Advanced Image Generation: Negative Prompt Optimization and Latent Classifier Guidance
arXiv:2607.14580v1 Announce Type: new Abstract: We present a novel system that integrates negative prompt optimization via a fine-tuned sequence-to-sequence LLM and latent-space classifier guidance to improve the quality of images generated by Stable Diffusion. Our approach automatically generates optimized negative prompts, and employs a CNN-RNN hybrid classifier to evaluate and guide diffusion steps, rolling back low-quality latent updates. Experimental results demonstrate that our dual-guidance framework reduces artifacts and improves semantic fidelity compared to baseline diffusion.
Thermodynamic theory of voting and EU elections
arXiv:2607.15119v1 Announce Type: cross Abstract: We introduce a thermodynamic theory of voting and show that it provides a good description of distribution of party votes in EU elections. The theory traces parallels between system energies of coupled nonlinear oscillators and party vote fractions. Such a classical system evolution is characterized by the conservation of total energy and probability norm that leads to the Rayleigh-Jeans (RJ) thermalization and condensation at low energy states. A similar thermalization also describes the wealth inequality in society. This feature belongs to the phenomena of constraint driven condensation known in statistical mechanics. We show that the RJ theory well depicts the Lorenz and Pareto curves obtained from the EU vote results. The theory also recovers the dispersion of votes between candidates of first round presidential elections in France.
Wichmann-Kroll Correction to the Interelectronic Interaction in He- and Li-Like Ions
arXiv:2607.12168v2 Announce Type: replace Abstract: We present a theoretical study of the higher-order QED contribution to the interelectronic interaction in He- and Li-like ions, where a virtual electron-positron loop is inserted into the photon line of the one-photon exchange diagram. Our approach is based on the Dirac-Coulomb Green's function and accounts for the interaction of the virtual $e^+e^-$ pair with the electric field of the nucleus to all orders in $\alpha Z$, with $\alpha$ being the fine-structure constant and $Z$ the atomic charge number. We show that the numerical convergence of the involved integrals can be significantly improved by explicitly subtracting the non-gauge-invariant spurious contributions from the integrands. We present improved numerical values for this contribution to the Lamb shift over a wide range of nuclear charge numbers $Z$. Our calculations agree well with previous results by Artemyev and co-workers [Phys. Rev. A 56, 3529 (1997); Phys. Rev. A 60, 45 (1999)] for He-like ions, but we find a discrepancy in the Li-like case. Moreover, we calculate the finite nuclear size correction to this diagram, which can reduce its size by more than 5% for heavy ions. The improved QED calculations not only decrease the uncertainty of theoretical predictions for the interelectronic interaction in few-electron ions but the methods could also be used in the future to improve calculations of closely related one-electron two-loop QED diagrams.
Self-Aware Recursively Self-Improving Agents for Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture
arXiv:2607.12254v2 Announce Type: replace Abstract: Large language model (LLM) agents can plan, use tools, maintain memory, and execute long-horizon tasks. This paper proposes Self-Aware Recursively Self-Improving (SARSI) agents: governed agents that maintain a persistent self-model of identity, goals, capabilities, limitations, uncertainty, relationships, history, and developmental change, and use that model to guide and evaluate recursive improvement. Self-awareness is defined functionally and does not imply subjective experience or phenomenal consciousness. We pair SARSI agents with personal singularity, a bounded human-AI co-development objective in which an agent ecosystem helps a user approach an expanding, user-defined feasible capability frontier. Each agent has a goal contract, bounded scope, validated tool registry, tool tests, end-to-end benchmarks, owner-controlled autonomy, routing, memory, self-model, and improvement policy. A scope router assigns every accepted task to one accountable primary agent and transfers out-of-scope work through structured handoffs. A user-facing Auto-Index selects interactive, hybrid, autonomous, or scheduled behavior without overriding external permissions. The architecture combines a planner-executor-verifier loop, an evidence-gated improvement loop, an external governance plane, decentralized lineages, an owner-directed agent foundry, and a Personal Singularity OS coordinating working, computational-imaging, work-process-learning, and personal-learning agents. We formalize functional self-awareness, scope, routing, improvement acceptance, bounded goal evolution, tool-first execution, and human capability transfer, and provide safety invariants, benchmark design, and a staged implementation roadmap. This is a position and systems-design paper, not evidence that consciousness, unrestricted recursive self-improvement, or personal singularity has been achieved.
VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation
arXiv:2607.12356v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have emerged as a powerful end-to-end paradigm for robotic manipulation by mapping language instructions and 2D visual inputs directly to actions. However, these models lack an explicit, scene-level 3D representation, limiting their ability to reason over spatial layouts and geometric constraints. While recent efforts incorporate explicit 3D cues, such as depth maps or point clouds, to improve geometric awareness, they primarily capture low-level structures and lack high-level semantic grounding in 3D space. In human cognition, interaction with the physical world relies on a 3D semantic cognitive map - an internal mental model that integrates spatial layouts with semantic context to enable persistent, viewpoint-invariant reasoning. In light of this, we present VistaVLA, a novel two-stage framework that constructs a geometry- and semantics-aware 3D cognitive representation from 3D Gaussian primitives and grounds it as compact context tokens for VLA policy learning. Specifically, VistaVLA lifts multi-view vision-language features into 3D Gaussian primitives, forming geometry-anchored semantic tokens that align view-consistent spatial grounding with 2D visual feature spaces. To make this 3D representation computationally tractable for effective VLA control, we introduce Merge-then-Query (MtQ), a token summarization mechanism. MtQ compresses dense Gaussian primitives into a highly compact set of spatially informative tokens, achieving a 99% token reduction while preserving action-relevant 3D layouts and semantic context. Extensive evaluations in both simulated and real-world environments demonstrate the effectiveness of VistaVLA. Notably, in real-world scenarios, VistaVLA improves success rates by 22.8% across seven real-world tasks and by 30.0% over the VLA-Adapter baseline on challenging out-of-distribution tasks.
Still image and spatial-temporal tomato data enabling detection, segmentation, tracking, and video-instance segmentation using strong and weak labels
arXiv:2607.14934v1 Announce Type: new Abstract: In this manuscript we release two datasets for visual sensing of tomato plants grown in commercial-like settings and acquired using a robot. The first is BUTom21 which consists of still images and manual annotations. The second is BUTom-ST21 which consists of video-based data and semi-automated annotations through AI-based methods, referred to as pseudo-labels. In both cases, we provide pixel-level labels for the ripeness of the fruit. The aim is to provide the research community a challenging set of real-world imagery to explore methods to sense and estimate the state of tomato plants and their fruit, which is an important horticultural crop. Importantly, the spatial-temporal dataset provides individual fruit count and ripeness information enabling researchers to push the boundaries of field-based phenotyping.
Integration Matters: Rollout-Based Training for Constrained Diffusion Models
arXiv:2607.14398v1 Announce Type: new Abstract: Constrained generative models aim to produce samples that satisfy complex feasibility constraints while remaining faithful to the data distribution. Existing constrained generation methods typically enforce constraints either through training-time optimization or sampling-time correction. Training-time optimization approaches optimize on states induced by the training distribution, which can differ substantially from those encountered during sampling. Sampling-time correction methods instead modify the sampling process at inference, introducing distribution shift and requiring expensive tuning, particularly for few-step sampling. We propose a fine-tuning framework that incorporates constraint guidance obtained through online rollout into the training process, which aligns training with sampling by differentiating through the fixed noise schedule used to numerically integrate the denoising process. This exposes the model to violations that arise along the denoising trajectory and aligns diffusion learning with the sampling process. Experiments across multiple tasks show that our method improves constraint satisfaction while maintaining competitive sampling quality compared to prior methods.
Dynamic mode decomposition for detecting oscillatory transient activity via sparsity and smoothness regularization
arXiv:2508.10266v4 Announce Type: replace Abstract: Dynamic mode decomposition (DMD) is a data-driven modal decomposition technique that extracts coherent spatio-temporal structures from high-dimensional time-series data. By decomposing the dynamics into a set of modes, each associated with a single frequency and a growth rate, DMD enables a natural modal decomposition and dimensionality reduction of complex dynamical systems. However, when DMD is applied to transient dynamics, even if a large number of modes are used, it remains difficult to interpret how these modes contribute to the transient behavior. In this study, we propose a simple extension of DMD that facilitates extraction of oscillatory transient activity by introducing time-varying amplitudes for the DMD modes based on sparsity and smoothness regularization. This approach enables identification of dynamically significant modes and extraction of their transient activities, providing a more interpretable representation of non-steady dynamics. We illustrate the validity of the proposed method using a simple example and then apply it to fluid flow data of a laminar airfoil wake exhibiting transient behavior. We demonstrate that it can capture the temporal structure of mode activations that are not accessible with the standard DMD method.
Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech
arXiv:2607.08208v2 Announce Type: replace Abstract: This paper describes our self-designed system for Task 1 of the MLC-SLM 2026 Challenge for multilingual two-speaker conversational speech. The system combines a modular speaker diarization front end with a challenge-adapted Qwen3-ASR-1.7B recognizer. The diarization front end performs voice activity detection, subsegment generation, CAMPPlus speaker embedding extraction, two-speaker spectral clustering, and RTTM-based audio segmentation. The resulting speaker-attributed segments are grouped by language or region and decoded by the adapted ASR model. For ASR adaptation, we first perform supervised full fine-tuning on the official training data, then apply LoRA fine-tuning with synthetic speech generated by a three-pipeline TTS-based synthetic speech augmentation framework, and finally refine the model using GRPO reinforcement learning with rewards based on WER/CER and penalties for hallucination, repetition, and length deviation. On the official development set, the full system achieves an average tcpMER of 23.70, reducing the error rate by 6.83 absolute points relative to the released Qwen-ASR-1.7B performance. On the final evaluation set, the system achieves an average tcpMER of 17.97. Ablation results show that supervised fine-tuning provides the largest gain, while synthetic-speech LoRA adaptation and reinforcement learning further improve robustness.
Aurora DSQL: Scalable, Multi-Region OLTP
arXiv:2607.13276v2 Announce Type: replace Abstract: Aurora DSQL is a serverless SQL database designed for cloud-scale transaction processing with multi-region active-active capabilities. Built on a disaggregated architecture, DSQL separates compute, storage, and transaction coordination into independent, horizontally scalable services. Query processors run in Firecracker MicroVMs executing PostgreSQL-compatible SQL without local state. The system uses multiversion concurrency control with precision timestamps for coordination-free reads and optimistic concurrency control for writes, deferring coordination to commit time through distributed adjudicators and the Journal replication system. This minimizes cross-region latency by requiring coordination only during commits, not individual statements. DSQL enables elastic scaling from zero to millions of transactions per second while providing strong consistency, ACID transactions, and continuous availability during availability zone or region failures.
RACiMo: Red Ambiental Ciudadana de Monitoreo: A Student-Centred Citizen Science Network for Environmental Monitoring, Data Literacy, and Climate Awareness in Colombia
arXiv:2607.14805v1 Announce Type: new Abstract: This paper traces the evolution of Red Ambiental Ciudadana de Monitoreo (RACiMo), a citizen-science and environmental-education initiative developed by Universidad Industrial de Santander in Colombia. RACiMo has evolved through successive editions from a school-based open-hardware programme into a regional environmental monitoring network that integrates meteorological and air-quality data, open-data practices, and student-centred research. The project's central objective is to train citizens, especially secondary-school students, teachers, and university mentors, to record, curate, analyse, interpret, and communicate climate and air-quality data relevant to their local contexts. RACiMo has involved approximately 515 students from urban, metropolitan, rural, and p\'aramo communities in Santander, Colombia. It has progressed from Arduino-Raspberry Pi stations assembled through a Do-It-With-Others approach to low-maintenance commercial instruments and a multi-municipality network of professional weather and air-quality stations. Its pedagogy has shifted from learning by building sensors to learning by analysing real environmental datasets through Python, Jupyter notebooks, visualisation tools, and student-led research projects. The current RACiMo-Orqu\'ideas implementation expands the network across five municipalities, emphasises local environmental issues such as freight traffic and industrial activity, and provides public access to data through repositories and interactive tools. The paper discusses the trade-offs between openness, educational value, reliability, data quality, sustainability, and scalability. It argues that citizen environmental monitoring can strengthen climate awareness and adaptation when local communities are trained not only to collect data, but also to understand and use it critically.
Emergent Region-Level Facial Correspondence in Frozen Vision Foundation Models
arXiv:2607.14423v1 Announce Type: new Abstract: Frozen self-supervised vision models can align parts of generic objects, but it remains unclear whether this correspondence extends to human faces, where global layout is shared while identity-specific appearance varies sharply. We test whether frozen DINOv3 features define a region-level facial coordinate system: a feature space in which eyes, brows, nose, mouth, skin, and hair remain distinguishable across people and across time without face-specific training. Using DINOv3 ViT-L/16 patch embeddings and FaRL only as a face-part labeling interface, we evaluate cross-identity nearest-neighbor matching and temporal label propagation on 200 CelebDF-v2 real videos. DINOv3 achieves 83.0% region-level semantic accuracy under unconstrained cross-identity matching, compared with a 23.0% area-weighted random baseline, and 95.5% temporal tracking accuracy without a learned temporal module. A no-FaRL control collapses to 0.9%, showing that FaRL supplies semantic initialization while DINOv3 supplies dense spatial correspondence. The strongest correspondence appears at an intermediate layer: block 18 gives a 4.93x same-region versus cross-region discrimination ratio, compared with 1.48x at the final block. Against CLIP ViT-L/14, DINOv3 shows only a small aggregate advantage but a +16.8 pp gain on anatomical regions, indicating that image-level contrastive supervision captures coarse facial layout but not fine-grained anatomical identity. These results establish frozen DINOv3 as a strong zero-shot representation for region-level facial correspondence and identify intermediate self-supervised features as the most useful layer for dense face analysis.
An unfitted boundary algebraic equation method with static-dynamic reduction for evolving implicit geometries
arXiv:2607.14808v1 Announce Type: new Abstract: Repeated elliptic solves on domains with evolving boundaries arise in moving-interface simulation, design, and reactive navigation. Even when a fixed Cartesian grid avoids remeshing, rebuilding all boundary interactions for every configuration can limit the efficiency of repeated solves. We develop a static--dynamic boundary reduction for an unfitted lattice Green's function method on prescribed moving planar domains. Like boundary integral and boundary element methods, the formulation reduces the problem to boundary-supported unknowns through a Green representation. Its construction, however, reverses the usual order: the Cartesian operator is discretized before the Green representation is formed, rather than representing the continuous problem first and then discretizing the boundary. This discretize-then-represent viewpoint avoids boundary meshes and singular quadrature. The method also separates interactions associated with stationary geometry from those affected by motion, reuses the invariant part throughout a simulation, and updates only couplings involving the changing boundary. Boundary conditions are imposed at true interface intersections, lattice-kernel data are reused, and the interior field is reconstructed by a fast sine-transform solver. The principal contribution is an implemented and validated update strategy for translating, deforming, appearing, and topology-changing obstacles.
A Deterministic Binary Fingerprinting Framework with Zero-Trained Feature Extraction for Sparse Count Matrices
arXiv:2607.14596v1 Announce Type: new Abstract: Sparse count matrices from single-cell transcriptomes to k-mer profiles and document-term frequencies are conventionally analyzed via PCA-reduced graph clustering or iterative optimization in continuous embedding spaces. We introduce MMTB, a deterministic non-learned binary representation framework that requires no label supervision, model fitting, or gradient-based optimization. Column-wise Min-Max normalization followed by fixed cutoffs maps each sample to a thermometer fingerprint whose Hamming distances show empirical correspondence with normalized L1 distances, with Pearson correlation approximately 0.92 on single-cell RNA-seq pairs. In a favorable three-cell-line mixture, the 3-threshold fingerprint achieves NMI of 0.99 at 188 bytes per cell. Under a fair Hamming nearest-neighbor graph plus Leiden readout, MMTB approaches PCA plus Leiden on this coarse task. On challenging tissue-like annotations, continuous pipelines often lead; PBMC Seurat NMI is 0.39 for MMTB versus 0.49 for Scanpy, underscoring that MMTB is suited for coarse-grained separation and resource-constrained deployments rather than fine-grained subtype discovery or as a general replacement for continuous embeddings. Relative to dense float32 representations, MMTB fingerprints reduce memory by approximately 10-fold while providing fixed-width Hamming-indexable codes. PCA30 embeddings and sparse CSR may be smaller; we do not claim universal compression. A label-free suitability score is provided as a deployment guideline, not a performance predictor.
An LLM-Based Automatic Sportscast Solution for Robot Soccer Matches
arXiv:2607.14809v1 Announce Type: new Abstract: RoboCup has always been a scenario to develop systems that solve real-world problems. Driven by the main goal of playing against the 2050 FIFA World Cup champions, the RoboCup Soccer leagues need to constantly measure how the research community is progressing. Computing visual statistics from match videos is a crucial way to track this evolution. To address this challenge, this paper introduces a fully autonomous, real-time sports commentator for RoboCup matches. By bridging the gap between raw kinematic tracking and natural language generation, our neuro-symbolic architecture extracts precise statistics from video streams and turns them into fluent, hallucination-free narration. The proposed system is capable of generating statistics and commentary both during live match streaming and in post-game analysis, easily adapting to the new dynamism of the league where different humanoid robots of different sizes share the field. Supplemental materials are available at https://lab-rococo-sapienza.github.io/MARIO/
Skeleton: Visual Authoring of Non-visual Data Experiences
arXiv:2607.14579v1 Announce Type: new Abstract: When sighted practitioners author accessible data visualizations, they build navigation structures (the nodes, edges, and input bindings that govern how assistive technologies traverse an interface) entirely in code, with no visual representation. Without a representation to react to, practitioners cannot develop judgment about what makes navigation good or bad, and the quality ceiling of non-visual experiences is set by the absence of a feedback loop. We address this problem through longitudinal co-design with practitioners across cartography, design systems, and open-source visualization, and make three contributions. First, we introduce an Inspector that renders navigation graphs as interactive node-link diagrams, and a Dimensions API that expresses navigation in terms of data dimensions rather than explicit graph construction. Second we present Skeleton, a direct-manipulation authoring environment in which the properties of an accessible navigation structure are translated into visual representations authors can observe and manipulate. Key techniques include a dual-view editor that simultaneously shows the system's navigation model and the end user's spatial experience, a scaffolding engine that automates spatial node placement by repurposing a visualization rendering pipeline, a live label-template editor with real-time screen-reader-output preview, and a testing mode that makes traversal sequence visually trackable. Third, we evaluate Skeleton through an in-situ study with 8 practitioners across visualization design, engineering, and research. Making navigation structure visible changed how practitioners engaged with accessible design: they reconsidered the architecture of their own visualizations, attended to a broader range of input modalities, and shifted from treating accessibility as a compliance task to treating it as a design problem. (abstract shortened for arxiv)
Interleaved Noise Injection Improves Clean, Corrupted, and OOD Performance
arXiv:2607.14466v1 Announce Type: new Abstract: Noise injection is a well-known technique in stochastic optimization. We report its surprising effectiveness with an interleaved (on-off-on-off...) rather than the usual monotonic decay schedule. We present a theoretical analysis of noise injection, which confirms that corruption by impulse noise approximates a Jacobian regularization, whereas Gaussian noise acts as a curvature penalty. This regularization behavior has been invoked to explain why noise injection increases model robustness. But the interleaved nature of our proposed schedule produces superior results even for the optimization objective: mixing phases of noisy data permits the optimizer to escape local minima and increase exploration without the risk of catastrophically forgetting the important features from the clean data. To stabilize this training scheme against the rapid changes of the loss when switching between clean and noisy data, we introduce a gradient-norm stabilization technique that scales noisy updates based on clean gradient magnitudes. We compare this method with other common augmentation methods and find substantial improvements in corruption tolerance and robustness to real-world distribution shifts on CIFAR-100-C, ImageNet-C, and ImageNet-R for ResNet and ViT architectures, with the best results being achieved by stacking our method on top of other augmentations. Through saliency and attention maps we show that the effect of interleaved noise injection stems from penalizing the failure modes encouraged by the inductive bias of the models: impulse noise works against the locality bias of convolutional (ResNet) architectures, and Gaussian noise reduces the tendency of attention-based models to pick up large-scale spurious features. Interleaved noise injection is therefore an effective tool to improve the test performance on clean, noisy, and out-of-distribution data at essentially zero computational cost.
Workload-Driven Optimization for On-Device Real-Time Subtitle Translation
arXiv:2607.09957v2 Announce Type: replace Abstract: This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-one inference, low latency, and privacy constraints. These conditions limit the value of optimizations designed for long-context or high-throughput language-model serving. Starting from LMT-60-0.6B, preliminary profiling suggests that vocabulary projection becomes a more important decode-time cost after GGUF quantization reduces the relative cost of Transformer blocks. We replace the original 151k-token vocabulary with a 64k-token subtitle-domain tokenizer, migrate the embedding space, and adapt the model through embedding calibration followed by full supervised fine-tuning. On an OpenSubtitles2024 test set, LocalSubs achieves a 59.2% tie-excluded win rate against Google Translate under GPT-4o pairwise judging. Performance is strongest on short cues and declines as cue length increases. In a separate preliminary Apple M2 Metal profiling run, LocalSubs shows a 1.63x speedup over a 151k-vocabulary baseline. The code is available on https://github.com/aiden1020/localsubs .
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
arXiv:2607.13278v2 Announce Type: replace Abstract: Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However, these models are typically specialized and do not generalize well to mixed or intermediate audio types. In this work, we adapt a diffusion-based model originally designed for multi-instrument music synthesis to voice conversion, covering both speech and singing within a unified framework. Specifically, we extend musical note-based conditioning to include phonetic posteriorgrams (PPGs) and pitch contours, and reinterpret timbre conditioning as speaker or singer identity via feature-wise linear modulation. Experiments show that the adapted model matches or surpasses a dedicated voice conversion system in terms of naturalness and performer similarity, while maintaining accurate pitch control across speech and singing. At the same time, we observe limitations in phonetic fidelity and a degradation in vocal quality when incorporating instrumental training data. Furthermore, we demonstrate that off-the-shelf feature extractors provide effective conditioning signals, enabling large-scale self-supervised training without manual annotations. These results highlight the potential of cross-domain model transfer towards unified audio generation systems capable of handling speech, singing, and music. Qualitative samples can be found on our project page: https://benadar293.github.io/voice-conversion