arXiv:2506.18414v3 Announce Type: replace Abstract: Melanoma is a highly aggressive skin cancer, making early and accurate diagnosis critical. While deep learning excels in skin lesion classification, standard ``black-box" models struggle to explain diagnostic uncertainty, limiting clinical trust. This work introduces a hybrid framework combining a class-aware adversarial Variational Autoencoder and an XGBoost classifier, transcending simple binary classification by leveraging a generative latent space for interpretable decision support. Guided by adversarial training, the model learns the visual characteristics of skin lesions and projects them into a continuous latent space, ensuring that similar images are grouped closely together. Trained on this latent space, the XGBoost classifier achieves a robust AUC of 0.868, competing closely with state-of-the-art models. For borderline cases, the framework enables clinicians to leverage the latent topology through Content-Based Image Retrieval. This provides a dual benefit: it allows the clinician to visually compare an ambiguous lesion against biopsy-confirmed precedents and acts as an early warning sign since a borderline classification can indicate that a lesion shares features of both nevi and melanomas, potentially requiring close monitoring. Our approach translates algorithmic hesitation into transparent, evidence-based visual support, bridging the gap between predictive performance and clinical trust.
Science Journals
arXiv:2508.07952v2 Announce Type: replace Abstract: Clustering algorithms often assume all features contribute equally to the data structure, an assumption that usually fails in high-dimensional or noisy settings. Feature weighting methods can address this, but most require additional parameter tuning. We propose SHARK (Shapley Reweighted $k$-means), a feature-weighted clustering algorithm motivated by the use of Shapley values from cooperative game theory to quantify feature relevance, which requires no additional parameters beyond those in $k$-means. We prove that the $k$-means objective can be decomposed into a sum of per-feature Shapley values, providing an axiomatic foundation for unsupervised feature relevance and reducing Shapley computation from exponential to polynomial time. SHARK iteratively re-weights features by the inverse of their Shapley contribution, emphasising informative dimensions and down-weighting irrelevant ones, and is equivalent to replacing the arithmetic mean of feature dispersions with their harmonic mean. Experiments on synthetic and real-world data sets show that SHARK consistently matches or outperforms existing methods, achieving superior robustness and accuracy, particularly in scenarios where noise may be present. Software: https://github.com/rickfawley/SHARK.
arXiv:2606.25934v1 Announce Type: new Abstract: A proof of optimal-order $H^1$-norm error estimates is given for $A$-stable backward difference full discretizations (of order 1 and 2) of Willmore flow for closed two-dimensional surfaces. The numerical method discretizes a coupled system of evolution equations by evolving surface finite elements of polynomial degree at least two in space and backward difference method of order 1 or 2 in time. The convergence analysis is based on a stability analysis, based on energy estimates exploiting the anti-symmetric structure of the second-order system, in combination with Dahlquist's $G$-stability and the multiplier techniques of Nevanlinna and Odeh, with a new upper bound in the spirit of Dahlquist. Numerical experiments illustrate and complement the theoretical results.
arXiv:2508.20440v5 Announce Type: replace Abstract: When solving time-dependent partial differential equations (PDEs), traditional physics-informed neural networks (PINNs) may encounter several challenges. In particular, standard PINNs do not explicitly account for the temporal evolution order of time-dependent problems during training, which may affect the quality of temporal evolution in time-dependent PDEs. In addition, a single neural network may face difficulties in simultaneously representing different physical behaviors across multiple regions of the computational domain.To address these issues, we propose a domain-decomposition-based causal PINNs (DDC-PINNs) framework. The term causal refers to the fact that the temporal evolution is performed sequentially through classical ordinary differential equation (ODE) integration, thereby respecting the natural temporal ordering of time-dependent PDEs. The proposed framework enhances spatial approximation through domain decomposition and employs a sequential temporal-evolution strategy for time-dependent problems.Within this framework, an approximate solution is first obtained using domain-decomposition PINNs. Subsequently, the time-derivative term in the original PDE is retained, while the remaining solution-dependent terms are replaced by the obtained approximation, thereby transforming the original PDE into an auxiliary ODE system. Classical numerical methods for ODEs are then employed to perform temporal evolution without repeated neural-network optimization. As a result, DDC-PINNs decouples spatial approximation from temporal evolution while preserving the temporal evolution order through sequential ODE integration.Numerical experiments on several benchmark problems demonstrate the effectiveness of the proposed framework and provide proof-of-concept validation of the DDC-PINNs methodology.
arXiv:2606.25472v1 Announce Type: new Abstract: This study investigates propeller hydrodynamics at intermediate Reynolds numbers (Re), crucial for small-scale robotic systems but still uncharted. Experiments on a propeller-driven underwater vehicle and numerical simulations reveal thrust reversal--a phenomenon where clockwise propeller rotation leads to backward motion--in the approximate range 1.3 < Re < 150 under specific conditions. Notably, counterclockwise rotation consistently results in backward motion. Simulations reveal that this behavior arises when centrifugal suction, an inward force along the axis caused by radial outward flow from the propeller's rotation, dominates over fluid backward acceleration, the primary thrust mechanism at high Re. These findings provide critical insights into the unique dynamics of the intermediate Re regime and inform the design of efficient propulsion systems for miniature aquatic robots.
arXiv:2509.15900v2 Announce Type: replace Abstract: This work aims to predict blood flow with non-Newtonian viscosity in stenosed arteries using convolutional neural network (CNN) surrogate models. An alternating Schwarz domain decomposition method is proposed which uses CNN-based subdomain solvers. A universal subdomain solver (USDS) is trained on a single, fixed geometry and then applied for each subdomain solve in the Schwarz method. Results for two-dimensional stenotic arteries of varying shape and length for different inflow conditions are presented and statistically evaluated. One key finding, when using a limited amount of training data, is that incorporating a physics-aware constraint, as, in our case, flow rate conservation, into the USDS improves the prediction accuracy and convergence behavior of the Schwarz method compared to a purely data-driven USDS. As the USDS is a data-driven, inexact subdomain solver, admissible parameter ranges for the geometry and inflow configurations must be defined and tested.
arXiv:2509.19049v3 Announce Type: replace Abstract: Aggregation of Asian student data can reinforce the model minority myth by obscuring educational disparities among Asian student subgroups. This study investigated variation in conceptual physics knowledge across Asian racial and ethnic subgroups using data from the LASSO platform, analyzing responses from 16,810 students enrolled in 493 introductory calculus-based physics courses across 64 U.S. institutions. We applied Multilevel Analysis of Individual Heterogeneity and Discriminatory Accuracy to examine predicted pre- and posttest performance on the Force Concept Inventory and Force and Motion Conceptual Evaluation. The findings revealed performance differences among 19 Asian subgroups that the pan-Asian strata (the single aggregated Asian group) concealed. Subgroup predicted means spanned 15.8 percentage points on the pretest and 15.4 percentage points on the posttest. The lowest-performing subgroup's posttest mean was roughly equal to the highest-performing subgroup's pretest mean, indicating a performance gap of about a full semester of instruction. Mean absolute error between the pan-Asian strata and the 19-subgroup estimates was 3.9 percentage points at pretest and 4.0 percentage points at posttest, equivalent to approximately 4-5 weeks of learning in a 16-week course. These findings demonstrate that fine-grained identity data collection can support identifying disparities that common aggregation practices conceal.
arXiv:2510.04584v2 Announce Type: replace Abstract: Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. However, subtle changes, such as shifting the order of choices, result in substantially different results. Existing MCQA frameworks do not account for this variability and report a single accuracy number per benchmark or category. We dive into the MCQA evaluation framework and conduct a systematic study spanning three benchmarks (MMAU, MMAR and MMSU) and four models: Audio Flamingo 2, Audio Flamingo 3, Qwen2.5-Omni-7B-Instruct, and Kimi-Audio-7B-Instruct. Our findings indicate that models are sensitive not only to the ordering of choices, but also to the paraphrasing of the question and the choices. Finally, we propose a simpler evaluation protocol and metric that account for subtle variations and provide a more detailed evaluation report of LALMs within the MCQA framework.
arXiv:2510.04773v2 Announce Type: replace Abstract: As Large Language Models (LLMs) demonstrate remarkable capabilities learned from vast corpora, concerns regarding data privacy and safety are receiving increasing attention. LLM unlearning, which aims to remove the influence of specific data while preserving overall model utility, is becoming an important research area. One of the mainstream unlearning classes is optimization-based methods, which achieve forgetting directly through fine-tuning, exemplified by Negative Preference Optimization (NPO). However, NPO's effectiveness is limited by its inherent lack of explicit positive preference signals. Attempts to introduce such signals by constructing preferred responses often necessitate domain-specific knowledge or well-designed prompts, fundamentally restricting their generalizability. In this paper, we shift the focus to the distribution-level, directly targeting the next-token probability distribution instead of entire responses, and derive a novel unlearning algorithm termed \textbf{Di}stribution \textbf{P}reference \textbf{O}ptimization (DiPO). We show that the requisite preference distribution pairs for DiPO, which are distributions over the model's output tokens, can be constructed by selectively amplifying or suppressing the model's high-confidence output logits, thereby effectively overcoming NPO's limitations. We theoretically prove the consistency of DiPO's loss function with the desired unlearning direction. Extensive experiments demonstrate that DiPO achieves a strong trade-off between model utility and forget quality. Notably, DiPO attains the highest forget quality on the TOFU benchmark, and maintains leading scalability and sustainability in utility preservation on the MUSE benchmark.
arXiv:2606.25206v1 Announce Type: new Abstract: Long-term robot deployment requires a compact and scalable memory that preserves fine-grained visual semantics, grounds observations in space and time, and enables efficient storage and retrieval. In this paper, we propose RAVEN, an agentic memory system for long-horizon robotic question answering and navigation. RAVEN stores visual embeddings with pose and time in a vector database, and grounds retrieval in a spatial map to answer queries and navigate to goals. By operating directly on visual embeddings, RAVEN avoids lossy image-to-text captioning and enables accurate semantic, spatial, and temporal retrieval at scale. Across several simulated and real-world video question-answering benchmarks, RAVEN consistently surpasses caption-based memory systems and matches frontier VLMs on long-horizon tasks at 10$\times$ lower retrieval cost. Finally, we instantiate RAVEN on a Unitree Go1 robot for the task of long-horizon navigation for natural language goal-reaching, and show successful deployment over several large indoor environments.
arXiv:2606.24891v1 Announce Type: new Abstract: Ontologies enable scalable energy services in buildings by supporting interoperability and automation. Project Haystack is a building ontology that is widely adopted due to its flexible, tag-based semantic model, openness, and extensibility, but suffers from ambiguous tag usage and limited automated validation. Although Project Haystack is formally open, its reliance on custom file formats and domain-specific languages that originate from the Haxall ecosystem creates a de facto barrier to integration. In this paper, we address these limitations by introducing a Python-based toolchain for Haystack. We present (i) a parser for Haystack definition files (Trio file format), and (ii) a code generator that derives Pydantic models and JSON Schema definitions from these parsed specifications. The resulting models enable static type checking and enable structural validation of Haystack grids within Python, as well as schema-based validation of JSON representations outside the Python ecosystem. All tools, generated models, and schemas are released publicly under an open-source license, with the goal of strengthening the Haystack ecosystem and opening a practical pathway beyond its current technical boundaries.
arXiv:2510.07094v2 Announce Type: replace Abstract: This work focuses on sampling strategies of configuration variations for generating robust universal locomotion policies for quadrupedal robots. We investigate the effects of sampling physical robot parameters and joint proportional-derivative gains to enable training a single reinforcement learning policy that generalizes to multiple parameter configurations. Three fundamental joint gain sampling strategies are compared: parameter sampling with (1) linear and polynomial function mappings of mass-to-gains, (2) performance-based adaptive filtering, and (3) uniform random sampling. We improve the robustness of the policy by biasing the configurations using nominal priors and reference models. All training was conducted using the RaiSim simulation environment, tested in simulation on a range of diverse quadrupeds, and zero-shot deployed onto hardware using the ANYmal quadruped robot. Compared to multiple baseline implementations, our results demonstrate the need for significant joint controller gains randomization for robust closing of the sim-to-real gap.
arXiv:2510.21700v2 Announce Type: replace Abstract: In this paper we construct distance sketches for intersection graphs of arbitrary path-connected regions in the plane (known as the string graphs) in the constant and $1+\varepsilon$ distortion regimes. Furthermore, the distance sketches themselves are planar graphs. First, we show that every unweighted string graph $G$ has an $O(1)$-distortion planar emulator: that is, there exists an edge-weighted planar graph $H$ containing every vertex in $G$, such that every pair of vertices $(u,v)$ satisfies $\delta_G(u,v) \le \delta_H(u,v) \le O(1) \cdot \delta_G(u,v)$. Furthermore, we show that for any constant $\varepsilon > 0$, there is an edge-weighted planar graph $H'$ such that every pair of vertices $(u,v)$ satisfies $\delta_G(u,v) \le \delta_{H'}(u,v) \le (1+\varepsilon) \cdot \delta_G(u,v) + O(\varepsilon^{-4}\textrm{poly}\log n)$. No previous constructions of sparse distance sketches were known even for intersection graphs of simple shapes like axis-parallel rectangles or fat convex polygons. As applications, we construct the first $(1+\varepsilon, +O(1))$ mixed-distortion tree cover and distance oracle for arbitrary string graphs, as well as the first additive $+(\varepsilon\Delta+O(1))$-distortion embedding of string graphs $G$ with diameter $\Delta$ into graphs of constant treewidth $O(\varepsilon^{-4})$.
arXiv:2511.22411v2 Announce Type: replace Abstract: 3D head stylization enables expressive reimagining of human faces for creative visual experiences in digital media. Existing 3D-aware methods often require computationally intensive optimization or per-style fine-tuning, limiting flexibility and user control. To overcome these challenges, we introduce StyleFusion360, a diffusion-based framework for multi-view consistent, identity-preserving 3D head stylization from a single style reference image, without per-style training. Our approach enhances the Style Fusion Attention mechanism with a style-conditioned key modulation mechanism that aligns content and style representations for fine-grained and controllable stylization. We further provide a user-controllable slider for adjusting stylization intensity. In addition, StyleFusion360 supports local multi-edit stylization, enabling targeted edits such as modifying hair or eyes independently. Extensive experiments on FFHQ and RenderMe360 demonstrate that StyleFusion360 produces high-quality, controllable, and visually compelling stylizations, outperforming state-of-the-art GAN- and diffusion-based methods across diverse style domains.
arXiv:2512.01759v3 Announce Type: replace Abstract: We investigate the potential of weights to serve as effective representations, focusing on neural fields. Our key insight is that constraining the optimization space through a pre-trained base model and low-rank adaptation (LoRA) can induce structure in weight space. Across reconstruction, generation, and analysis tasks on 2D and 3D data, we find that multiplicative LoRA weights achieve high representation quality while exhibiting distinctiveness and semantic structure. When used with latent diffusion models, multiplicative LoRA weights enable higher-quality generation than existing weight-space methods.
arXiv:2512.22487v2 Announce Type: replace Abstract: The design of Korean constituency treebanks raises a central representational question concerning the choice of terminal units. Although Korean words are morphologically complex, treating morphemes as constituency terminals can obscure the distinction between word-internal morphology and phrase-level syntactic structure, and can create mismatches with eojeol-based dependency resources. This paper argues for an eojeol-based constituency representation, with morphological segmentation and fine-grained POS information encoded in a separate, non-constituent layer. A comparative analysis shows that, under explicit normalization assumptions, the Sejong, Penn Korean, and KAIST treebanks can be compared over a shared eojeol-based constituency backbone. Building on this result, we outline an eojeol-based annotation scheme that preserves interpretable constituency, supports cross-treebank comparison and constituency-dependency alignment, and provides a surface-form terminal layer for future end-to-end Korean constituency parsing.
arXiv:2605.12767v2 Announce Type: replace Abstract: Radioactive molecules provide a powerful new platform in the search for new physics at energy scales complementary to high-energy particle colliders. By combining enhancements from nuclear properties with the sensitivity and control offered by molecular structure, experiments with radioactive molecules offer great reach in the search for new physics beyond the Standard Model. Rapid progress in this field is being driven by advances in the production and control of radioactive molecules, alongside the development of new experimental tools and theoretical techniques. In this Perspective, we discuss the current status and future prospects of this rapidly developing, interdisciplinary field at the intersection of nuclear physics, atomic and molecular physics, and particle physics.
arXiv:2601.09117v3 Announce Type: replace Abstract: Generative AI systems increasingly enable the production of highly realistic synthetic media. Civitai, a popular community-driven platform for AI-generated content, operates a monetized feature called Bounties, which allows users to commission the generation of content in exchange for payment. To examine how this mechanism is used and what content it incentivizes, we conduct a longitudinal analysis of all publicly available bounty requests collected over a 14-month period following the platform's launch. We find that the bounty marketplace is dominated by tools that let users steer AI models toward content they were not trained to generate. At the same time, requests for content that is "Not Safe For Work" are widespread and have increased steadily over time, now comprising a majority of all bounties. Participation in bounty creation is uneven, with 20% of requesters accounting for roughly half of requests. Requests for "deepfake" - media depicting identifiable real individuals - exhibit a higher concentration than other types of bounties. A nontrivial subset of these requests involves explicit deepfakes despite platform policies prohibiting such content. These bounties disproportionately target female celebrities, revealing a pronounced gender asymmetry in social harm. Together, these findings show how monetized, community-driven generative AI platforms can produce gendered harms, raising questions about consent, governance, and enforcement.
arXiv:2606.25221v1 Announce Type: new Abstract: Optical 3D scanning systems allow the acquisition of accurate models of patient anatomy, suitable for use in the design of simple 3D-printable patient-matched medical devices with 3D modelling software. This study developed and demonstrated the use of superficial brachytherapy surface mold design workflow that utilizes data from optical 3D surface scanning and enables a commercial brachytherapy treatment planning system to be used for catheter positioning and dose optimization steps. Synthetic CT images were generated from 14 optically scanned anatomical models of human participants. Models and skin textures ac-quired from the optical scans were imported into Autodesk Meshmixer, where the treatment area was delineated, and treatment and device volumes produced. 3D Slicer was used to convert the body, treatment and device volumes to DICOM CT and RTSTRUCT data. The synthetic CT data and contoured volumes were imported into Varian Eclipse, where catheters were designed, and dwell positions and times optimised for dose coverage of the treatment volume. The lack of in-ternal anatomy did not compromise dose calculations, due to clinical use of a TG43 based algorithm. Once 3D printed, molds can be imaged in-situ during CT simulation, and reconstructed, for clinical dose calculation and plan approval.
arXiv:2601.15477v2 Announce Type: replace Abstract: Biological membranes are dynamic surfaces whose shape and function are critically influenced by protein inclusions (PIs). While membrane deformations induced by PIs have been extensively studied in the small-deformation regime, a variety of processes involves strong membrane deformations. We investigate the interaction between lipid membranes and PIs in the large deformation (LD) regime, with the finite-element method. We develop an approximate analytical solution that captures key features of the LD regime. We show that the force exerted by the membrane on a PI displays a non-monotonic behavior with respect to the PI vertical displacement. The qualitative features of this force appear to be independent of the protein geometry. For two interacting PIs, the membrane-mediated potential exhibits sub-power-law decay with inter-protein distance, reflecting the complex nature of the elastic medium. The interaction potential shows that conical PIs with identical and opposite orientations repel and attract, respectively, confirming the analogy between PI orientation and electric charge, in the LD regime. In the presence of membrane flows, we identify a characteristic velocity that separates two regimes in which bending rigidity and viscous effects dominate, respectively, implying the onset of flow-induced deformations above such velocity threshold. Overall, our results provide quantitative predictions for membrane-protein systems in biologically relevant scenarios involving LDs, with implications for protein sorting, clustering, and membrane trafficking.
arXiv:2601.16096v3 Announce Type: replace Abstract: We introduce Neural Particle Automata (NPA), a Lagrangian generalization of Neural Cellular Automata (NCA) from static lattices to dynamic particle systems. Unlike classical Eulerian NCA where cells are pinned to pixels or voxels, NPA model each cell as a particle with a continuous position and internal state, both updated by a shared, learnable neural rule. This particle-based formulation yields clear individuation of cells, allows heterogeneous dynamics, and concentrates computation only on regions where activity is present. At the same time, particle systems pose challenges: neighborhoods are dynamic, and a naive implementation of local interactions scale quadratically with the number of particles. We address these challenges by replacing grid-based neighborhood perception with differentiable Smoothed Particle Hydrodynamics (SPH) operators backed by memory-efficient, CUDA-accelerated kernels, enabling scalable end-to-end training. Across tasks including morphogenesis, point-cloud classification, and particle-based texture synthesis, we show that NPA retain key NCA behaviors such as robustness and self-regeneration, while enabling new behaviors specific to particle systems. Together, these results position NPA as a compact neural model for learning self-organizing particle dynamics.
arXiv:2601.17037v2 Announce Type: replace Abstract: We investigate visual reasoning limitations of both multimodal large language models (MLLMs) and image generation models (IGMs) by creating a novel benchmark to systematically compare failure modes across image-to-text and text-to-image tasks, enabling cross-modal evaluation of visual understanding. Despite rapid growth in machine learning, vision language models (VLMs) still fail to understand basic visual concepts such as object orientation, quantity, and spatial relationships, which highlights gaps in elementary visual reasoning. By adapting MMVP benchmark questions into explicit and implicit prompts, we create \textit{AMVICC}, a novel benchmark for profiling failure modes across various modalities. After testing 11 MLLMs and 3 IGMs in 9 categories of visual reasoning, our results show that failure modes are often shared between models and modalities. However, certain failures are model-specific and modality-specific, and this can potentially be attributed to various factors. IGMs consistently struggle to manipulate specific visual components in response to prompts, especially in explicit prompts, suggesting poor control over fine-grained visual attributes. Our findings apply most directly to the evaluation of existing state-of-the-art models on structured visual reasoning tasks. This work lays the foundation for future cross-modal alignment studies, offering a framework to probe whether image generation and visual interpretation failures stem from shared limitations. These insights can guide future improvements in unified vision-language modeling.
arXiv:2601.21416v2 Announce Type: replace Abstract: The generalization capabilities of robotic manipulation policies are heavily influenced by the choice of visual representations. Existing approaches typically rely on representations extracted from pre-trained encoders, using two dominant types of features: global features, which summarize an entire image via a single pooled vector, and dense features, which preserve a patch-wise embedding from the final encoder layer. While widely used, both feature types mix task-relevant and irrelevant information, leading to poor generalization under distribution shifts, such as changes in lighting, textures, or the presence of distractors. In this work, we explore an intermediate structured alternative: Slot-Based Object-Centric Representations (SBOCR), which group dense features into a finite set of object-like entities. This representation permits to naturally reduce the noise provided to the robotic manipulation policy while keeping enough information to efficiently perform the task. We benchmark a range of global and dense representations against intermediate slot-based representations, across a suite of simulated and real-world manipulation tasks ranging from simple to complex. We evaluate their generalization under diverse visual conditions, including changes in lighting, texture, and the presence of distractors. Our findings reveal that SBOCR-based policies outperform dense and global representation-based policies in generalization settings, even without task-specific pretraining. These insights suggest that SBOCR is a promising direction for designing visual systems that generalize effectively in dynamic, real-world robotic environments.
arXiv:2601.22615v3 Announce Type: replace Abstract: Streaming recurrent models enable efficient 3D reconstruction by maintaining persistent state representations. However, they suffer from catastrophic forgetting over long sequences due to balancing historical information with new observations. Recent methods alleviate this by deriving adaptive signals from the attention perspective, but they operate on single dimensions without considering temporal and spatial consistency. To this end, we propose a training-free framework termed TTSA3R that leverages both temporal state evolution and spatial observation quality for adaptive state updates in 3D reconstruction. In particular, we devise a Temporal Adaptive Update Module that regulates update magnitude by analyzing temporal state evolution patterns. Then, a Spatial Contextual Update Module is introduced to localize spatial regions that require updates through observation-state alignment and scene dynamics. These complementary signals are finally fused to determine the state updating strategies. Extensive experiments show that TTSA3R achieves competitive performance on standard short-sequence benchmarks and provides substantially stronger robustness on extended sequences. On NRGBD, as sequences extend from 50 to 250 frames, TTSA3R exhibits only a 1.33x error increase, compared with over 4x degradation for CUT3R. This highlights the practical value of temporal-spatial adaptive updates for long-term reconstruction stability. Our code is available at https://github.com/anonus2357/ttsa3r.
arXiv:2602.01903v3 Announce Type: replace Abstract: This work studies online episodic tabular Markov decision processes (MDPs) with known transitions and develops best-of-both-worlds algorithms that achieve refined data-dependent regret bounds in the adversarial regime and variance-dependent regret bounds in the stochastic regime. We quantify MDP complexity using a first-order quantity and several new data-dependent measures for the adversarial regime, including a second-order quantity and a path-length measure, as well as variance-based measures for the stochastic regime. To adapt to these measures, we develop algorithms based on global optimization and policy optimization, both built on optimistic follow-the-regularized-leader with log-barrier regularization. For global optimization, our algorithms achieve first-order, second-order, and path-length regret bounds in the adversarial regime, and in the stochastic regime, they achieve a variance-aware gap-independent bound and a variance-aware gap-dependent bound that is polylogarithmic in the number of episodes. For policy optimization, our algorithms achieve the same data- and variance-dependent adaptivity, up to a factor of the episode horizon, by exploiting a new optimistic $Q$-function estimator. Finally, we establish regret lower bounds in terms of data-dependent complexity measures for the adversarial regime and a variance measure for the stochastic regime, implying that the regret upper bounds achieved by the global-optimization approach are nearly optimal.