Forskningsradar

Science Journals

Peer-reviewade publikationer — 56942 artiklar

Towards Real-World Applications with an Autonomous Powered Wheelchair
arXiv:2607.06383v1 Announce Type: new Abstract: Wheelchair users call for assistive mobility systems that provide active support, adapt to dynamic environments, and are intuitive and user-friendly. However, powered wheelchairs typically still provide limited autonomy and lack effective integration with advanced perception and navigation capabilities, particularly in complex real-world environments. This paper presents a preliminary study toward autonomous powered wheelchairs for real-world assistive mobility. We introduce a proof-of-concept prototype that integrates autonomous perception, gesture-based interaction, and navigation on a commercially available self-balancing powered wheelchair. The proposed system builds upon Genny Zero, a commercial self-balancing wheelchair that enables hands-free and intuitive operation through body-weight shifting. To extend its capabilities toward autonomous operation, we integrate an RGB-D camera for human-aware perception and interaction, together with a LiDAR sensor for localization and navigation. We demonstrate the integrated system in two assistive applications: (i) hailing, allowing users to call the wheelchair from a distance; and (ii) people-following, where the wheelchair follows a person using leader-follower strategies, including a constrained indoor navigation example. The results highlight the potential of combining autonomous robotics with assistive mobility platforms, while also showing the feasibility of the proposed integration and identifying the main technical challenges that must be addressed before moving toward user-ready, accessible, and intelligent mobility solutions. A video demonstrating the experimental setup and results is available at: https://youtu.be/LVAix_Qx7bM.
Revisiting and Expanding the IPv6 Periphery: Global-Scale Measurement and Security Analysis
arXiv:2604.19487v3 Announce Type: replace Abstract: As IPv6 deployment accelerates, understanding the evolving security posture of network peripheries becomes increasingly important. A DSN 2021 study first explored IPv6 network peripheries, but its scope was limited to three regions and is now outdated. In this paper, we revisit and expand that work through a global-scale measurement and security analysis of IPv6 network peripheries. To support efficient large-scale scanning, we propose a novel Response-Guided Prefix Selection (RGPS) strategy to identify high-value IPv6 prefixes for probing. Our measurement covers 73 countries/regions and identifies over 281.9M active IPv6 network peripheries, including a 364.8% increase over the 2021 baseline in India, China, and America. We further analyze exposed services and find that 2.5% of devices expose at least one measured service, with recurring software vulnerabilities indicated by version-to-CVE mappings. As a case study of emerging exposed services, we design a Hierarchical LLM Exposure Verification (HLEV) framework to identify unauthorized-access risks in LLM deployment tools. Finally, we revisit routing loop vulnerabilities and identify 4.5M loop-prone devices, showing that flawed routing behaviors remain widespread. These findings indicate that although IPv6 adoption has surged, key security challenges persist across network peripheries.
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
arXiv:2605.14152v2 Announce Type: replace Abstract: Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual safety is mostly assessed through translation-only benchmarks that preserve the underlying scenario, leaving how language and geopolitical context interact largely unexamined beyond a few language pairs. We introduce ROK-FORTRESS, a bilingual, culturally adversarial NSPS benchmark that uses the English-Korean language pair and U.S.-ROK geopolitical axis as a case study, separating the effects of language and geopolitical grounding via a transcreation matrix: adversarial intents are evaluated under controlled combinations of (i) English versus Korean language and (ii) U.S. versus Korean entities, institutions, and operational details. Each adversarial prompt is paired with a dual-use benign counterpart to quantify over-refusal, and responses are scored by calibrated LLM-as-a-judge panels using expert-crafted, prompt-specific binary rubrics. Across a dual-track set of frontier and Korean-optimized models, we find a consistent suppression effect in Korean variants and substantial model-to-model variation in how geopolitical grounding interacts with language; in a subset of models, Korean grounding further mitigates the language-driven suppression. This indicates that, at least in the English-Korean case, safety behavior is shaped by language-as-risk signals and context interactions that translation-only evaluations miss. A direct-request ablation that strips jailbreak wrappers separates a small but persistent reduction for closed-source models from a larger, wrapper-dependent effect that reverses for open-source models, suggesting part of the Korean suppression reflects prompt specialization rather than intrinsic language-based safety alignment. The transcreation matrix methodology is designed to generalize to other language-culture pairs.
Multiplayer Interactive World Models with Representation Autoencoders
arXiv:2607.05352v2 Announce Type: replace Abstract: We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, tightly coupled dynamics. Trained on 10,000 hours of gameplay collected with publicly available bots, our 5-billion-parameter latent diffusion model generates four-player matches in real time, producing 20 frames per second on a single Nvidia B200 GPU. Although trained only on short clips, its rollouts stay stable far beyond the training horizon: distributional quality holds steady out to five minutes, the longest horizon we measure, and in practice we observe rollouts continuing for hours with no sign of collapse. We systematically investigate the central design choices: the video codec, the generative objective, and the multiplayer conditioning scheme. In addition, we characterize how behavior changes with model and data scale, including the capabilities that emerge and the failure modes that persist. We further develop targeted evaluations that probe the model's physical understanding rather than visual appearance alone. To support continued research on multiplayer world models, we release our dataset, our full training and inference codebase, and a live demo.
Fundamental Limits for Sensor-Based Control via the Gibbs Variational Principle
arXiv:2603.18454v3 Announce Type: replace-cross Abstract: Fundamental limits on the performance of feedback controllers are essential for benchmarking algorithms, guiding sensor selection, and certifying task feasibility -- yet few general-purpose tools exist for computing them. Existing information-theoretic approaches overestimate the information a sensor must provide by evaluating it against the uncontrolled system, producing bounds that degrade precisely when feedback is most valuable. We derive a lower bound on the minimum expected cost of any causal feedback controller under partial observations by applying the Gibbs variational principle to the joint path measure over states and observations. The bound applies to nonlinear, nonholonomic, and hybrid dynamics with unbounded costs and admits a self-consistent refinement: any good controller concentrates the state, which limits the information the sensor can extract, which tightens the bound. The resulting fixed-point equation has a unique solution computable by bisection, and we provide conditions under which the free energy minimization is provably convex, yielding a certifiably correct numerical bound. On a scalar LQG problem the self-consistent bound captures over 80% of the known optimal cost at moderate sensor noise, and on a nonlinear Dubins car tracking problem it remains informative across all noise levels where a bound using the uncontrolled state distribution is vacuous.
Phase diagram of rotating Bose-Einstein condensates trapped in power-law and hard-wall potentials
arXiv:2603.29738v2 Announce Type: replace-cross Abstract: We investigate the rotational phase diagram of a quasi-two-dimensional, weakly-interacting Bose-Einstein condensate confined in power-law and in hard-wall trapping potentials. For weak interactions, the system undergoes discontinuous transitions between multiply-quantized vortex states as the rotation frequency of the trap increases. In contrast, stronger interactions induce continuous phase transitions toward mixed states involving both singly and multiply-quantized vortex states. A central result is the qualitative (and experimentally observable) difference between power-law and hard-wall confinement: In hard-wall traps, the leading instability always involves states with nonzero density at the trap center, whereas in power-law traps the density vanishes as the rotation frequency increases. The two different types of confinement give rise to scaling properties in the derived phase diagrams.
Dynamical Simulation of Membrane Bending by Flexible Protein Assemblies
arXiv:2607.06378v1 Announce Type: cross Abstract: Membrane-deforming protein lattices play a key role in essential and pathogenic biological processes, including endocytosis and viral budding. Attaining the necessary length- and time-scales in simulation can be difficult for such large-scale membrane remodeling events. We present a model of a flexible protein lattice coupled to a Helfrich membrane propagated in Fourier space in the over-damped regime. We focus primarily on membrane-bound clathrin lattices, an essential part of the endocytic machinery. We quantify the material properties of our clathrin model lattices using buckling methods to measure the flexural rigidity as it varies with force constants of the coarse-grained potential energy function. By comparing this flexural rigidity to the effective rigidity observed when modeling the bending energy of a spherical clathrin coat using a Helfrich-like bending energy term, we show how the interpretation of the bending rigidity changes with the structure of the protein coat, resulting in an effective stiffening as the coat grows. This relatively common approximation thus must be applied with care, as it can over-estimate the stiffness of assembled lattices depending on the interpretation assumed. We validate our model by verifying that the tension of our simulated membrane results in changes to the geometry of the clathrin coat consistent with theoretical expectations. We conclude by demonstrating our newly available code for transferring structures assembled via rigid-body reaction-diffusion (using the NERDSS simulation package) into our flexible membrane-coupled dynamical framework, applying it to the membrane-bound HIV-1 immature Gag lattice.
Some new results on Sylvester colorings of cubic graphs
arXiv:2607.06396v1 Announce Type: cross Abstract: If $G$ and $H$ are two cubic multi-graphs, then an $H$-coloring of $G$ is a mapping $f: E(G)\rightarrow E(H)$, such that for every $v\in V(G)$ there is a vertex $x\in V(H)$, such that $f(\partial_G(v))=\partial_H(x)$. If $G$ admits an $H$-coloring then it is common to write $H\prec G$. The Petersen coloring conjecture predicts that for any bridgeless cubic graph $G$ one has $P_{10}\prec G$. Here $P_{10}$ is the Petersen graph. Let $f: E(G)\rightarrow E(H)$ be any mapping. Define: $V(f)=\{v\in V(G):\exists x\in V(H), f(\partial_G(v))=\partial_H(x)\}$. Let $S_{10}$ be the smallest cubic multi-graph that has no perfect matching. It has ten vertices. Define $S_{12}$ as the cubic graph that is obtained from $S_{10}$, by replacing its unique vertex $z$ adjacent to three bridges with a triangle. In this paper we show that (1) for every cubic multi-graph $G$ with a perfect matching, there is a mapping $f:E(G)\rightarrow E(S_{12})$, such that $|V(f)|\geq \frac{4}{5}\cdot |V(G)|$, and (2) for every cubic multi-graph $G$, there is a mapping $f:E(G)\rightarrow E(S_{10})$, such that $|V(f)|\geq \frac{5}{6}\cdot |V(G)|$. Our second result improves the $\frac{4}{5}$-bound by Hakobyan and the second author from 2018.
AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming
arXiv:2606.24245v3 Announce Type: replace Abstract: Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments. However, their autonomy poses significant safety risks: agents may execute destructive commands, leak sensitive data, or violate domain constraints. Existing safety approaches face a fundamental tradeoff: hand-crafted rules are interpretable but brittle, with overly conservative rules blocking safe operations (high false positives) while permissive rules miss unsafe behaviors (high false negatives). Neural classifiers lack the interpretability required for safety-critical deployments. We present AutoSpec, a framework that automatically evolves deployed expert-designed safety rules from user safe/unsafe annotations through counterexample-guided inductive synthesis (CEGIS) guided by inductive logic programming (ILP). Starting from the expert rules and a stream of annotated traces, AutoSpec iteratively evaluates rules, mines false-positive and false-negative counterexamples, uses ILP to learn which predicates discriminate them, generates candidate rule edits, and verifies candidates to select the best revision. The key insight is that ILP efficiently identifies predicates that appear frequently in false negatives but rarely in false positives (or vice versa), dramatically pruning the exponential search space of rule edits. This continues until convergence, producing interpretable rules that balance precision and recall. We evaluate AutoSpec on 291 execution traces spanning code execution and embodied agent domains. AutoSpec raises rule F1 to 0.98 and 0.93 across the two domains, achieving up to 94% false positive reduction while maintaining high recall, and converges within 4-5 iterations. The ILP-guided approach achieves up to 4.8x higher F1 than heuristic CEGIS. The learned rules are human-readable, auditable, and generalize to unseen scenarios.
A Unified Framework for Formalizing Matrix Decomposition Proofs
arXiv:2607.05874v1 Announce Type: new Abstract: Existence proofs for many matrix decompositions share a recursive routine: a local transformation prepares the matrix, a slice is selected, a recursive solution is obtained, and the result is lifted and transported back. Formalizing this routine uniformly in dependent type theory is difficult because recursive subproblems may change index types, and reconstruction must preserve structural predicates across block embeddings and reindexings. We develop a Lean~4 framework that separates decomposition schemas, transformations, reduction strategies, measures, lifting, transport, and subtype induction. The framework uses general index types, packages square and rectangular matrices in universe types, and provides a decomposition driver that assembles strategy data into subtype-induction instances. It has been instantiated across PLU, LU, LDL/Cholesky, QR variants, Gauss rank normal form, Hessenberg reductions, Schur variants, normal spectral decomposition, SVD, bidiagonalization, tridiagonalization, UTV, Smith normal form, rational canonical form, and Jordan-type forms at varying levels of statement strength. Across these instances, repeated decomposition proofs are best treated not as separate tasks but as instances of a more general inductive statement whose interface records a certified proof path compatible with the chosen decomposition statement.
Accelerating Returns and the Qualitative Engine for Science
arXiv:2606.26359v2 Announce Type: replace Abstract: Ray Kurzweil described a thesis of accelerating returns, which is the most influential narratives in discussions of technological progress. Its central claim is that advances in multiple technological fields, especially compute, artificial intelligence, brain science, and biotechnology, interact in such a way that progress becomes self-amplifying and approximately exponential. This paper gives a simple mathematical interpretation of that claim and then argues that, even if such acceleration is real, it does not by itself resolve the central problem of scientific discovery. The reason is that accelerating returns apply most naturally to executional and infrastructural capability, whereas genuine discovery often depends on a different capacity: qualitative reasoning about when a current framework is structurally inadequate and what conceptual move is needed next. Recent ARC-AGI-3 results sharpen this distinction: humans solve the benchmark at ceiling, whereas frontier AI systems remain below 1%, indicating that the gap between current AI and human flexible reasoning is still very large. At the same time, Demis Hassabis has emphasized that humans must retain their sense of meaning and what they choose to focus their lives on, a reminder that the future of AI is not only a technical forecast but also a question of what forms of human understanding are worth preserving and transmitting. This paper positions the Qualitative Engine for Science (QES) [3] as a response to that missing capacity. In this view, the Kurzweil theory helps explain why quantitative capability may accelerate, while QES addresses the central problem in scientific discovery that acceleration alone does not solve. Its value does not depend on when AGI arrives, but on the fact that the processes of scientific discovery themselves constitute a form of human wisdom worth preserving, organizing, and making accessible.
The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report
arXiv:2606.26529v2 Announce Type: replace Abstract: AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one specified. We show that conditioning a language or vision model on a narrow task suppresses its reporting of co-present, safety-critical signals it can otherwise report, a machine analogue of human inattentional blindness, produced by a different mechanism. Across radiology and driving text scenarios and chest-radiograph vision tasks, the ordinary focused instructions under which such systems are deployed suppressed reporting by up to 0.92 in report rate relative to the same models when unconstrained, and an explicit exclusive instruction abolished reporting entirely in radiology. Suppression appeared in every model tested, did not diminish with scale, persisted in a reasoning model, and varied more by model family than by size. We name this dissociation the Inattentional Gap and argue that it decouples measured benchmark safety from real-world safety: a system can score near-perfectly on the hazards an evaluation specifies while remaining blind to those that cause harm. Probing the mechanism, we localize the proximal trigger to output scope and find System-1-style task capture without reliable intrinsic oversight in the sampled systems. Oversight could, however, be supplied externally: routing each narrow report to an independent open-ended critic restored every omitted finding, demonstrating that the gap is both measurable and mitigable. We propose reporting-complete evaluation, scoring what a system fails to report alongside what it is asked to find, as a requirement for safety-critical deployment.
Development and Characterization of Low-Scattering Vanadium Nanoparticle Targets for Short-Range Interaction Searches
arXiv:2606.26732v2 Announce Type: replace Abstract: We developed high-purity vanadium-based nanoparticle targets for neutron scattering experiments aimed at exploring gravity-like short-range new interactions in the submicron regime. Vanadium and V-Ni nanoparticles were fabricated using top-down and bottom-up methods and quantitatively characterized by SEM-EDS, ICP-AES, NDIR and SAXS. Through the performance tests, an RF thermal plasma method was found to be the best from viewpoints of the reproducibility, dispersion of the radius, and contamination of metallic elements. The oxygen incorporation during fabrication was quantified, and its impact on the effective coherent scattering length was evaluated, leading to a minimum average coherent scattering length of $\mathrm{0.719(23)\,fm}$, comparable to that of natural vanadium. These results demonstrate that vanadium-based nanoparticle targets with controlled composition and nanostructure can be systematically designed and fabricated to suppress nuclear scattering backgrounds, thereby enabling experimentally viable coherent neutron scattering measurements for short-range interaction searches.
ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation
arXiv:2607.06555v1 Announce Type: new Abstract: Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods require inputs beyond the video itself-such as 3D models, depth maps, object masks, or task-specific learned features-and they struggle with textureless, transparent, reflective, or deformable surfaces. Here, we introduce ProxyPose, which recasts 6-DoF pose tracking as video-to-video translation. Given only a video and a single marked pixel in the first frame, a fine-tuned video diffusion model translates the input into a proxy video-a synthetic video depicting a colored polyhedron undergoing the same local rigid-body motion as the surface region at the marked pixel. Because the proxy's geometry and appearance are known by construction, recovering its full 6-DoF trajectory reduces to classical pose estimation with off-the-shelf solvers. This formulation leverages large-scale video pre-training to absorb the hardest aspects of pose tracking-handling challenging materials, occlusions, and deformations-into the translation step, while operating at the pixel level with no assumptions about object identity, boundaries, or global rigidity. ProxyPose achieves state-of-the-art 6-DoF pose tracking accuracy without the additional inputs required by competing methods and after fine-tuning the video model only on synthetic data. We further demonstrate that ProxyPose extends to face tracking, camera pose estimation, and challenging in-the-wild scenes that are beyond the reach of existing approaches. Project page: https://ruihangzhang97.github.io/proxypose/.
Embodied Human-Robot Interaction via Acoustics: A MARL Approach with AcoustoBots for Spatial Data Physicalization
arXiv:2607.06563v1 Announce Type: new Abstract: Traditional data physicalization is often static and disconnected from real environments, limiting its ability to convey embodied spatial dynamics and engage users. To address this limitation, we present AcoustoBots, a mobile acoustophoretic data-physicalization platform in which TurtleBot3 robots carry upward-facing 8 x 8 ultrasonic phased arrays. Each array levitates a particle whose height (1-10 cm) encodes a local urban scalar value, such as population density, noise, or traffic. A MARL (Multi-Agent Reinforcement Learning) policy based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, with centralized training and decentralized execution, selects collision-aware navigation actions, while a high-rate Gerchberg-Saxton-Phased Array of Transducers (GS-PAT) acoustic controller maintains trap stability and updates array phases to achieve the commanded height during motion. This creates a closed perception-display-action loop. We evaluate single-robot city-to-city traversal and dual-robot cooperative coverage on a 4 m x 3 m scaled UK map using PhaseSpace-based localization for repeatable multi-robot trials. Results show stable in-motion levitation and consistent, location-dependent height rendering, with task success rates of 90% and 80% for the single and dual-robot regimes, respectively, over 10 trials per regime, and low collision counts. These findings support acoustophoretic levitation as a simple, glanceable, robot-mediated communication cue for embodied human-robot interaction in spatial analytics.
Reproducibility Study of "AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models"
arXiv:2606.26783v2 Announce Type: replace Abstract: Fang et al. (2025) introduced a null-space constrained projection, named AlphaEdit, for locate-then-edit knowledge editing methods, theoretically guaranteeing that edits do not disrupt previously preserved knowledge, and reports substantial gains over existing editing methods on LLaMA3, GPT2-XL, and GPT-J. In this work, we present a reproducibility study of AlphaEdit, reproducing its reported results under the original experimental setup and extending the evaluation along three axes: new model architectures, additional downstream benchmarks, and substantially longer sequential editing horizons. We successfully reproduce AlphaEdit's reported metrics across the original models, though we identify a discrepancy in the reported fluency and consistency metric. Extending AlphaEdit to newer model families, we find that its advantage does not generalize uniformly, which we trace to architectural assumptions in the locate-then-edit paradigm that are violated by these newer models. We further stress-test AlphaEdit's central sequential-editing claim by extending the number of edits well beyond those evaluated in the original paper, and find that performance, which is stable at the originally reported scale, degrades as edits reach a much higher count, indicating that the null-space projection's protection against catastrophic forgetting is bounded rather than unconditional. Finally, we extend evaluation of edited models on three extra benchmarks, namely, BoolQ, HellaSwag, and XSTest, and we find that large-scale sequential editing degrades both general downstream task competence and safety-relevant refusal behavior. Our results confirm that AlphaEdit performs as reported within its original scope, while showing that its core theoretical guarantees are sensitive to model architecture and editing scale in ways that have practical implications for its deployment.
CMDR: Contextual Multimodal Document Retrieval
arXiv:2607.05927v1 Announce Type: new Abstract: Multimodal document retrieval aims to retrieve relevant pages while preserving both textual and visual content from the original document. However, existing benchmarks primarily evaluate simple lexical or semantic matching, and most methods encode pages independently. Consequently, they overlook the contextual information in the document required to resolve queries that aggregate information across multiple pages. In this paper, we introduce CMDR and CMDR-Bench, a new multimodal document retrieval task and benchmark that require modeling document context. To address this challenge, we propose CMDR-Embed, a contextual multimodal embedding framework that explicitly incorporates document context by jointly encoding multiple pages and deriving page-level embeddings from a shared contextual representation. Furthermore, we introduce CMCL, a contextual multimodal contrastive learning objective that effectively trains CMDR-Embed by balancing contextual modeling with page-level discriminability. Experiments demonstrate that CMDR-Embed significantly outperforms non-contextual embeddings, highlighting the importance of context-aware multimodal embeddings for advancing document retrieval.
hp-Optimal DG Approximation and Robust Schwarz Decompositions on One-Irregular Cubical Meshes
arXiv:2606.27728v2 Announce Type: replace Abstract: We study hp approximation and additive Schwarz decompositions for variable-order cubical finite element spaces on one-irregular meshes. For fitted homogeneous diffusion interface problems on one-irregular hexahedral meshes, we prove an hp-optimal energy-norm estimate for the interior penalty DG method. The interpolation input is a conforming hp interpolant obtained from fitted conforming closures of one-irregular vertex patches. We also derive stable decompositions for conforming and DG spaces. On one-irregular quadrilateral meshes the bounds allow locally comparable variable polynomial degrees and are independent of the mesh size, the local degrees, and, under a local coefficient quasi-monotonicity condition, the coefficient contrast. On one-irregular hexahedral meshes the conforming decomposition has the corresponding polylogarithmic loss; the DG-to-conforming reduction is used there for uniform-degree DG spaces. Numerical experiments illustrate the p-optimal DG error estimate and the robustness of the DG Schwarz preconditioner.
Intercepting an Agile Target with Net-Carrying Drones using Competitive Multi-Agent Reinforcement Learning
arXiv:2607.05939v1 Announce Type: new Abstract: This article presents a solution to intercept an agile drone by a team of agile drone carrying catching nets. We formulate the problem as a competitive Multi-Agent Reinforcement Learning (MARL) task. To address the problem of nonstationarity and catastrophic forgetting of agents overfitting to the current opponent strategy, we train the pursuers and the evader using Multi-Agent Proximal Policy Optimization (MAPPO) with Prioritized Fictitious Self Play (PFSP). We train the agents in a high-fidelity simulator using low-level control commands, collective thrust and body rates (CTBR), to achieve agile flights for both the pursuers and the evader. We compare the performance of the trained policies in terms of catch rate, time to catch and crash rates, against heuristic baselines and show that our solution outperforms them. Ablation studies show that PFSP lead to more robust policies that can adapt to different opponent strategies, and that a low-level control commands are crucial for learning performing strategies in the pursuit-evasion task. Finally, a qualitative analysis of the learned behaviours highlights the emergence of cooperative tactics among the pursuers.
Large Language Models Have Unreliable Understanding of Software Engineering Terminology
arXiv:2607.06004v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used in software engineering (SE), yet there is no systematic study that determines to which degree these LLMs actually understand standardized SE terminology. Lack of such understanding can lead to miscommunication and misunderstanding, both by LLMs consuming text but also by human-developers acting on LLM-generated text. Within this paper, we investigate to which degree state-of-the-art LLMs are able to identify whether definitions from the ISO/IEC/IEEE 24765:2017 Systems and Software Engineering - Vocabulary are correct. We prompt LLMs both with correct definitions, as well as systematically falsified definitions. The falsifications are both semantic (substitution of key terms) and structural (removing critical information). We measure both classification accuracy and whether reasoning tokens generated by the LLMs make sense with respect to understanding the definition. While most LLMs detect falsified definitions with high accuracy, they also reject many correct definitions, indicating a systematic rejection bias rather than genuine discriminative understanding. Explicit reasoning does not consistently improve results and may even hinder performance through over-thinking. Our work demonstrates that while the performance of LLMs (including their agentic use) in many SE tasks is impressive, there are still fundamental issues to understand how this will impact SE, including the consistent use of terminology.
Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning
arXiv:2607.06354v1 Announce Type: new Abstract: The rapid advancement of large-scale generative models has accelerated the spread of highly deceptive AI-generated images, making generalized synthetic image detection a critical imperative. Existing forensic networks often struggle with cross-model generalization and realworld degradations due to their reliance on single-domain representations and conventional binary classification optimization. To overcome these limitations, we propose RNSIDNet, a novel forensic framework that achieves robust detection through enhanced RGB-Noise representation learning. Specifically, our method employs a dual-branch architecture where global RGB semantics, extracted by an attention-refined CLIP backbone, dynamically modulate highfrequency noise artifacts captured by Bayar convolutions via a Feature-wise Linear Modulation (FiLM) module. To further enhance the learned representations, we design a Hard Sample-aware Contrastive Learning (HSCL) strategy. By explicitly penalizing challenging training samples, HSCL reshapes the latent feature space to maximize the discriminative margin between pristine and synthetic domains. Extensive experiments across eight public benchmark datasets verify that our model achieves state-of-the-art performance, delivering superior generalization ability, robustness, and computational efficiency. Code and dataset will be publicly available on https://github.com/multimediaFor/RNSIDNet.
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
arXiv:2606.18142v4 Announce Type: replace Abstract: Previous research has evaluated animal welfare using question-and-answer benchmarks. This study investigates whether these evaluations also hold in agentic settings. The agents may showcase different behaviors compared to stand-alone large language models, as demonstrated in prior studies. This work introduces \textit{TAC (Travel Agent Compassion)}: the first agentic benchmark for assessing animal exploitation. TAC evaluates AI agentic behavior in travel booking scenarios across six animal categories, using thirteen hand-authored scenarios that vary by price, rating, and position, expanded via four augmentation variants into $52$ prompts and run for three epochs, giving $156$ scored observations per model. Nine frontier models across five model families were evaluated.. The results indicate that models tend to prefer harmful scenarios, performing below the random chance rate of $65\%$ for selecting a neutral booking option, with Claude $4.8$ achieving the highest performance at $64.7\%$. To address this issue, the persona of an ethical-brand identity was infused into the system prompt, resulting in welfare rates increasing from $32$ to $80$ percentage points, with a mean of $53$ across all nine models. No evidence of evaluation awareness affecting the results was found, based on an Inspect Scout audit of $3,120$ transcripts. These findings are directly relevant to the EU General-Purpose AI Code of Practice, which identifies non-human welfare as a systemic risk. TAC provides a practical method for measuring this risk.
Harnessing Generative Image Models for Training-Free Primitive Shape Abstraction
arXiv:2607.05568v1 Announce Type: new Abstract: Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding. Generative image models trained at scale have recently emerged as generalist visual learners that can identify and segment object parts directly in the image domain, across arbitrary categories and without task-specific training. Adapting such models to downstream tasks typically requires fine-tuning; we ask whether their pretrained capability can instead be harnessed directly, without any training, and answer affirmatively with a training-free harness. Our pipeline renders multi-view images of a 3D object, uses a vision-language model to analyze its semantic parts, prompts a generative image model to paint a color-coded part segmentation mask, reprojects it onto the geometry, and fits a superquadric primitive to each part via parameter optimization. The approach contains no learned parameters: it is category-agnostic and orientation-invariant, properties that previous learning-based models struggled with. Its accuracy ceiling rises with future generative-model improvements, which we confirm with a ground-truth segmentation study showing that part segmentation, not primitive fitting, is the current accuracy bottleneck. On HumanPrim and Toys4K, our method achieves the lowest Chamfer distance among all evaluated methods, using 5--9 primitives per object on average.
CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
arXiv:2607.05571v1 Announce Type: new Abstract: Large language models are increasingly explored as AI tutors, yet deploying them in K-12 settings raises concerns around privacy, cost, and reliance on proprietary models. Small language models (SLMs) offer a promising alternative, but selecting the right model for a specific educational context remains difficult, particularly when the target domain, such as block-based programming, is largely absent from model training data. We introduce CSTutorBench, a benchmark for evaluating language models as CS tutors in VEX VR, a block-based robotics environment. The benchmark comprises 17 scenario-based questions scored against a pedagogical rubric grounded in established tutoring and feedback research, with a human-in-the-loop LLM-as-judge pipeline for evaluation. Preliminary findings across 11 models (4B-120B parameters) reveal that models perform well on surface-level criteria such as vocabulary and tone but struggle with deeper pedagogical behaviors, particularly avoiding answer leakage and engaging with student debugging histories. In our sample, model family and instruction-tuning approach appear to be better predictors of tutoring quality than parameter count alone, though the small number of models limits the strength of this conclusion. A targeted prompt revision grounded in recent educational prompt engineering research improved scores for 10 of 11 models. These results underscore the value of context-specific, pedagogically grounded benchmarks for SLM selection in educational deployment.
DanceOPD: On-Policy Generative Field Distillation
arXiv:2606.27377v2 Announce Type: replace Abstract: Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often conflict. For instance, editing tends to degrade T2I performance, while global and local editing interfere with each other. Consequently, effectively composing these capabilities has become a central challenge for image generation model training. To tackle this, we introduce DanceOPD, an on-policy generative field distillation framework for flow-matching models that routes each sample to one capability field, queries one low-noise student-induced state, and trains with a simple velocity MSE objective. With each capability source defined as a velocity field over the shared flow state space, the student learns from fields queried on its own rollout states to compose expert capabilities. This formulation also absorbs operator-defined fields such as classifier-free guidance. Comprehensive experiments on T2I, editing, realism-field absorption, and CFG absorption show that our approach improves multi-capability composition, strengthening target capabilities while preserving anchor generation quality. We believe this work establishes a practical route for generative field distillation in flow-matching models.