arXiv:2407.08094v3 Announce Type: replace-cross Abstract: We introduce the Binless Multidimensional Thermodynamic Integration (BMTI) method for nonparametric, robust, and data-efficient density estimation. BMTI estimates the logarithm of the density by initially computing log-density differences between neighbouring data points. Subsequently, such differences are integrated, weighted by their associated uncertainties, using a maximum-likelihood formulation. This procedure can be seen as an extension to a multidimensional setting of the thermodynamic integration, a technique developed in statistical physics. The method leverages the manifold hypothesis, estimating quantities within the intrinsic data manifold without defining an explicit coordinate map. It does not rely on any binning or space partitioning, but rather on the construction of a neighbourhood graph based on an adaptive bandwidth selection procedure. BMTI mitigates the limitations commonly associated with traditional nonparametric density estimators, effectively reconstructing smooth profiles even in high-dimensional embedding spaces. The method is tested on a variety of complex synthetic high-dimensional datasets, where it is shown to outperform traditional estimators, and is benchmarked on realistic datasets from the chemical physics literature.
Science Journals
arXiv:2602.22918v3 Announce Type: replace Abstract: Vision-language models (VLMs) can read text from images, but where does this optical character recognition (OCR) information enter the language processing stream? We investigate the OCR routing mechanism across three architecture families (Qwen3-VL, Phi-4, InternVL3.5) using causal interventions. By computing activation differences between original images and text-inpainted versions, we identify architecture-specific OCR bottlenecks whose dominant location depends on the vision-language integration strategy: DeepStack models (Qwen) show peak sensitivity at mid-depth (about 50%) for scene text, while single-stage projection models (Phi-4, InternVL) peak at early layers (6-25%), though the exact layer of maximum effect varies across datasets. The OCR signal is remarkably low-dimensional: PC1 captures up to 72.9% of variance. Crucially, principal component analysis (PCA) directions learned on one dataset transfer to others, demonstrating shared text-processing pathways. Surprisingly, in models with modular OCR circuits (notably Qwen3-VL-4B), OCR removal can improve counting performance (up to +6.9 percentage points), suggesting OCR interferes with other visual processing in sufficiently modular architectures.
arXiv:2604.20127v2 Announce Type: replace Abstract: Failures in complex systems often emerge through gradual degradation and the propagation of stress across interacting components rather than through isolated shocks. Democratic systems exhibit similar dynamics, where weakening institutions can trigger cascading deterioration in related institutional structures. Traditional reliability and survival models typically estimate failure risk based on the current system state but do not explicitly capture how degradation propagates through institutional networks over time. This paper introduces a trajectory-aware reliability modeling framework based on Dynamic Causal Neural Autoregression (DCNAR). The framework first estimates a causal interaction structure among institutional indicators and then models their joint temporal evolution to generate forward trajectories of system states. Failure risk is defined as the probability that predicted trajectories cross predefined degradation thresholds within a fixed horizon. Using longitudinal institutional indicators, we compare DCNAR-based trajectory risk models with discrete-time hazard and Cox proportional hazards models. Results show that trajectory-aware modeling consistently outperforms Cox models and improves risk prediction for several propagation-driven institutional failures. These findings highlight the importance of modeling dynamic system interactions for reliability analysis and early detection of systemic degradation.
arXiv:2602.08141v2 Announce Type: replace Abstract: Two parameters scale factor leading to bouncing cosmology is considered. We show that at some model parameters we obtain the deceleration parameter $q_0\approx -0.535$ at the current epoch which is in agreement with the Planck data. The equation for the transition point when the universe expands from acceleration to deceleration phases is obtained. We find the equation for the function $F(T)$ within the teleparallel gravity with torsion field $T$ which provides bouncing cosmology. For some parameters of the model the function $F(T)$ was computed. At the same time, in the framework of entropic cosmology, the associated entropy was obtained for particular model parameters. The equation of state for dark energy was obtained.
Reef restoration practitioners aim to preserve coral genetic diversity by protecting reefs and cultivating diverse genotypes in coral nurseries. However, cryptic genetic lineages in most corals complicate restoration strategies, as the role of between-lineage genetic divergence remains unclear regarding adaptation. In Montastraea cavernosa, researchers have identified cryptic lineages, some strongly segregated by depth. We conducted a ten-week reciprocal transplantation experiment using two cryptic lineages restricted to shallow water (<10m depth), with one lineage more common on nearshore reefs and the other on offshore reefs. We aimed to quantify lineage-specific responses to the environment that explain the genetic and ecological divergence between the two lineages. Surprisingly, the strongest response was not lineage-specific. Instead, both lineages exhibited strong and similar changes in growth and metabolomic profiles, depending on the transplantation habitat. These results suggest that cryptic lineages employ similar mechanisms of adaptation and acclimatization to environmental challenges, despite their genetic distinction.
LINE-1 retrotransposons are the only autonomous mobile elements still active in human genomes and remain a potent source of mutation, genome remodeling, and disease risk. However, young, full-length, potentially active copies (the elements most likely to shape present-day genomes) have been largely inaccessible to population-scale analysis because they are long, repetitive, and poorly resolved by short-read sequencing. Here, we use 47 phased long-read assemblies from the Human Pangenome Reference Consortium, representing 94 haplotypes, to build an allele-resolved view of recent human LINE-1 evolution. We identify 13,617 LINE-1 alleles with intact ORF1 and ORF2 across 683 unique insertion sites, revealing that every genome carries a distinct repertoire of potentially active source elements. These intact LINE-1 profiles recapitulate broad human population structure while exposing a large, rare, and population-enriched reservoir of mobile-element diversity missed by single-reference approaches. We also resolve a structurally variable chromosome 11 LINE-1 array, demonstrating that local duplication and rearrangement can amplify LINE-1 sequence independently of canonical retrotransposition. By comparing full-length LINE-1 sequences, we define activity signatures that separate ancient remnants from recently expanding lineages and uncover young LINE-1 groups whose activity is not fully explained by canonical subfamily labels. Sequence-network analyses further reveal a dynamic history of lineage turnover, in which successful source elements rise, seed new insertions, and are replaced by descendants marked by specific nucleotide changes. Together, these data transform human LINE-1s from a repetitive background into a resolved evolutionary system, linking insertion polymorphism, coding potential, population history, and recent retrotransposon adaptation. Our findings establish the human pangenome as a framework for discovering active source elements and for testing how mob…
Gross chromosomal rearrangements are a hallmark of many diseases and cancers. The study of their biogenesis and the mechanisms underlying their formation is greatly facilitated by the availability of genetic reporter assays in model organisms. We present here a novel GCR assay developed in fission yeast, a highly relevant model for understanding genome instability related to human biology. The reporter employs canavanine counter-selection to detect GCRs within a chromosomal context. Using this assay, we identified natural hotspots for GCRs, including inverted long terminal repeats (IR-LTRs). Structural analysis of GCR events showed that IR-LTR-induced GCRs mainly result in either terminal deletions with adjacent inverted duplications or repair via long-range break-induced replication (BIR). Deleting IR-LTRs reduces the GCR rate and reveals another hotspot driven by BIR between homeologous aldo/keto reductase genes on opposite arms of chromosome I. This is the first evidence that BIR can occur in S. pombe on long tracks reaching up to 600 kb. Besides highlighting genome rearrangement hotspots, the assay also identifies regulators of genome instability in fission yeast. Loss of Nup132, a component of the nuclear pore complex, increases IR-LTRs-induced GCRs, while the budding yeast homolog Nup133 has no effect on the stability of a structurally similar IR. In contrast, disrupting djc9, which encodes a conserved histone H3-H4 binding protein, decreases GCR rates. Overall, this sensitive GCR assay enables the identification of factors that control spontaneous and fragile motif-induced chromosomal instability, including those conserved in humans but lost through evolution in other organisms.
Maximum utilization of existing genetic variability in a breeding program depends on the efficient classification of the inbred lines into heterotic groups, particularly under stress conditions. This study applied practical breeding approaches to determine the mode of genetic inheritance for Striga resistance and proposes a weighted heterotic grouping method based on the general combining ability of multiple traits (WHGCAMT) and compares its effectiveness with other existing methods in classifying the inbred lines into heterotic groups in Striga-infested and optimum environments. Using Diallel design IV, 300 crosses were generated from 21 inbred lines and 4 standard testers. The crosses, along with six checks, were evaluated in an 18 x 17 alpha lattice design with two replications at two locations, in both artificial Striga-infested and Striga-free environments. The inbred lines were genotyped using DArTtag SNP markers. Phenotypic and genotypic data were analyzed using R. Analysis of variance revealed significant mean squares for hybrid, general combining ability (GCA), specific combining ability (SCA) and their interactions with environment. Significant positive and negative GCA and SCA effects were detected for grain yield and other measured traits. However, a larger proportion of additive gene action than non-additive gene action was observed for grain yield and most measured traits. The analysis of molecular variance also showed substantial genetic differences within and between clusters. Except for HSCA, the mean grain yield between the inter-group and intra-group hybrids was significant for each method. Pairwise comparison of the inter- and intra-group hybrids of all the methods showed significant differences between the WHGCAMT and all other methods in most cases. WHGCAMT consistently produced higher-yielding inter-group hybrids and lower-yielding intra-group hybrids, achieving breeding efficiency improvements of 55.8%, 4.3%, 15.7%, and 11.4% over the HSCA…
Accurate and reproducible assessment of foliar disease severity is essential for evaluating the performance of heterogeneous plant communities and understanding host-pathogen interactions. However, traditional visual scoring methods remain subjective, with limited precision, and difficult to scale in large phenotyping experiments. Here, we present a semi-automated image analysis workflow designed to quantify multiple foliar disease symptoms simultaneously on wheat flag leaves sampled from varietal mixtures. The workflow combines three methodological components: (i) a standardized protocol for leaf sampling and imaging, (ii) supervised machine learning segmentation using Random Forest implemented in Ilastik to classify multiple symptoms (powdery mildew and yellow rust), and (iii) a graphical user interface facilitating pipeline deployment by non-specialist operators. To evaluate the influence of image representation on classification performance, four color spaces (RGB, HSV, HLS, LAB) were systematically compared. The approach was validated using images of durum wheat flag leaves collected from a field experiment assessing eight-way varietal mixtures under natural fungal pressure. Cross-validation against manually annotated images demonstrated high segmentation accuracy across all symptom. Comparison among color spaces revealed only minor differences in performance. Overall, this workflow offers a cost-effective, annotation-efficient and reproducible alternative to deep learning approaches, leveraging open-source and actively maintained tools while requiring limited training data and enabling objective, reproducible and scalable disease phenotyping.
Biodiversity is commonly summarized by macroecological mean patterns, most prominently the species-area relationship (SAR) linking habitat area to expected species richness. Yet conservation, policy, and economic decisions increasingly require risk metrics: probabilities of rare but consequential biodiversity shortfalls, including local collapse. Such tail risks are central in finance and insurance but remain difficult to quantify in ecology because the data needed to estimate full richness distributions are rarely available at decision scales. Here we provide a mechanistic route from species-area relationships to biodiversity risk metrics. We show that when regional species abundances are well approximated by Fisher's log-series, a minimal immigration-extinction mechanism yields a closed-form stationary distribution for local richness whose structure tightly couples the mean SAR to richness variability and lower-tail probabilities. This coupling implies exact fluctuation-response identities and an explicit integral transform that reconstructs collapse probabilities and other tail risk measures directly from the mean SAR. These results define ecological analogues of financial risk metrics---such as collapse probability and lower-tail quantiles---without requiring direct estimation of the full richness distribution. Using high-resolution ForestGEO tree censuses spanning tropical, subtropical, and temperate forests, we find empirical support for these predictions across spatial scales. Together, our results show how widely measurable species-area relationships can be elevated from descriptive averages to operational tools for biodiversity risk assessment and reliability-based conservation planning.
Although resources are typically distributed continuously in space, species distributions often organize into discrete clusters. In his seminal paper, Turing demonstrated that such clusters can spontaneously arise in population densities, even when populations evolve in environments with continuously varying conditions. This phenomenon is known as Turing instability. In this work, we focus on two models grounded in population dynamics: a one-dimensional model based on the nonlocal Fisher-KPP equation, and a two-dimensional model involving an environmental gradient. We show that phenotypic clusters (sometimes referred to as "species") emerge in these models. We prove that they do not emerge because of Turing instability, but because of stochasticity, and that they disappear when stochasticity is reduced. First, for both models, we start our simulations with initial populations uniformly distributed in the state space. We show that phenotypic clusters quickly emerge and that the distances between them depend on the population size, that is, on the degree of stochasticity. Next, we start from already clearly defined phenotypic clusters. We identify three regimes in the connection between population size, the initial distances between clusters, and the distances between clusters at equilibrium. Last, on the two-dimensional model, we relax the hypothesis of complete clonality by varying the effective recombination rate, explore its effect on phenotypic clustering, and show that phenotypic clustering decays drastically with slight recombination.
Elucidating how habitat degradation facilitates extinction is critical for effective conservation efforts. Here, we propose integrating physiologically-structured population models into stochastic population viability analyses to assess how differing consequences of habitat degradation interact to drive extinction dynamics in a focal population. Using the isolated spectacled caiman Caiman crocodilus population/ecomorph from the Apaporis River as a case study, we find that threatening the resource base, which individuals increasingly rely upon, to outgrow vulnerable size ranges and mature accelerates extinction. We also found that when habitat degradation impacts both the primary adult and juvenile resource bases, this can have marked synergistic effects on threatening population viability. By contrast, destroying nesting sites has only a small effect on accelerating the impact of deteriorating prey availability. Through integrating community-level feedback between habitat degradation/change and population dynamics/structure, our approach provides a comparative framework for assessing the relative importance of distinct mechanisms through which habitat degradation ultimately drives extinction risk.
Wildlife vaccination could become a powerful strategy to mitigate disease-induced biodiversity losses, yet many vaccines for wildlife diseases provide only limited protection. Notably, tools to control the fungal pathogen Batrachochytrium dendrobatidis (Bd) are urgently needed for amphibian conservation. Laboratory experiments have demonstrated that prophylactic exposure to Bd metabolites increases host resistance, significantly reducing infection intensity in amphibians subsequently challenged with live Bd. Because Bd metabolites are non-infectious and applied topically, this treatment has potential to be administered to waterbodies to vaccinate and protect amphibians. We developed an agent-based model that indicated imperfect vaccination could reduce or amplify Bd infections at the population level, depending on degree of enhanced resistance or tolerance. Utilizing a Before-After-Control-Impact design with ten years of data, we conducted an ecosystem-level trial where we applied low levels of Bd metabolites or a sham control treatment to ponds in California and subsequently quantified Bd prevalence and infection intensity in metamorphosing Pacific chorus frogs (Pseudacris regilla). Unexpectedly, infection intensity was significantly greater in treated ponds relative to control ponds following metabolite addition. Additional model simulations indicated that this could occur via two mechanisms: (1) if treatment greatly increased tolerance alone or in combination with smaller increases in resistance, or (2) if a deleterious environmental interaction caused the treatment to increase susceptibility, rather than promote resistance. Future research is needed to determine whether tolerance or environmental factors drove heightened Bd infection intensities in this field trial to identify contexts in which this treatment can be used as a conservation tool.
Cliffs are environmentally extreme yet biodiversity-rich ecosystems that harbour specialist plants, many endemic and threatened. Plant persistence in these nutrient-poor substrates may depend on tightly linked soil- and root-associated microbial communities, which remain poorly understood. These interactions may become increasingly important with the global expansion of recreational climbing. While physical climbing impacts on vegetation are documented, potential chemical effects, from the use of climbing chalk (magnesium carbonate), on soil properties and plant-associated microbiota remain unknown. We sampled soils and roots beneath cliff-specialist and generalist plants, and unvegetated soils, across climbed and unclimbed routes in northern, central, and southern Spain. Soil physicochemical properties were quantified, fungal communities were characterized using ITS-metabarcoding, and structural equation modelling was used to disentangle direct and indirect effects. Climbing increased soil pH and altered soil chemical properties, driving shifts in fungal diversity and functional composition in soil and roots. The relative read abundance of root-associated symbiotrophic fungi declined, whereas arbuscular mycorrhizal fungi and pathogens increased in climbed cliffs. Overall effects were consistent, with cliff-specialist plants mediating nutrient and fungal shifts. Our findings show that climbing can reshape cliff soil chemistry and fungal communities, with potential cascading consequences for plant functional performance, nutrient dynamics, and ecosystem resilience.
Anxiety has been extensively studied in relation to memory, yet its dynamic association with spatial episodic memory in naturalistic clinical settings remains largely unexplored. We developed an anxiety-spatial-memory EMA protocol (asm-EMA) and deployed it in 30 epilepsy patients undergoing inpatient EEG monitoring, delivering combined momentary anxiety ratings and a validated spatial memory task pseudo-randomly every 90-150 minutes across multiple days. Subject-level asm-EMA means and session-to-session variability both correlated significantly with standard neuropsychological assessments, supporting the clinical validity of our design. Elevated within-person STAI-6 was selectively associated with faster retrieval responses, yet spatial memory accuracy was independent of all three anxiety measures, suggesting a shift in response strategy rather than memory impairment. Within-day anxiety showed short-term carryover between consecutive sessions, with little persistence beyond the next session. The asm-EMA protocol provides a feasible, autonomous framework for capturing moment-to-moment anxiety-memory dynamics in naturalistic settings.
RNA-seq experiments routinely identify thousands of differentially expressed genes, but translating these into biological insights and therapeutic hypotheses often requires integrating multiple tools. Existing web platforms such as iDEP, NetworkAnalyst, and GEPIA2 address individual steps, differential expression, network visualization, or TCGA queries, but lack a unified environment spanning raw data processing to clinical and pharmacological interpretation. TransXplorer (https://www.transxplorer.org) is a freely available web platform that addresses this limitation by integrating the complete RNA-seq analytical workflow. It supports processing from raw FASTQ files using HISAT2 or Salmon, as well as direct GEO dataset import with automated metadata handling. Differential expression analysis is implemented via DESeq2, edgeR, and limma-voom, followed by functional enrichment across more than 1,800 species using Bioconductor resources. Batch effects are automatically detected and corrected using a composite of PVCA, kBET, and Silhouette metrics without requiring predefined batch annotations. Downstream analyses include co-expression network construction (WGCNA), protein-protein interaction mapping (STRING), cell-type deconvolution, and transcription factor inference using integrated DoRothEA and TFLink resources. The platform further links gene signatures to drug candidates through DGIdb and OpenTargets and enables survival and tumour-normal comparisons across TCGA cohorts. Application to cardiac endothelial differentiation (GSE151427) and kidney renal papillary cell carcinoma (TCGA-KIRP) datasets demonstrates accurate batch correction, biologically consistent pathway enrichment, recovery of expected cell-type proportions, and identification of clinically relevant genes and drug candidates. TransXplorer is freely available without a login.
Chromosome-scale assemblies are increasingly available for non-model organisms, but functional annotation remains limited when deep evolutionary divergence erodes primary amino-acid sequence identity even though protein structural similarity can remain conserved. We present a hybrid annotation framework that decouples gene-model discovery from cross-species similarity assignment by combining Evo2-based ab initio prediction of exon-intron structures with ESM-2 protein-embedding-based structural similarity mapping. Applied to the sea lamprey, the framework derives high- or medium-confidence cross-species similarity assignments for 73,485 Evo2-derived translated protein models, including 35,395 high-confidence calls, and expands the deduplicated structural catalog to 31,286 loci, including 20,871 additions absent from the Ensembl baseline. A joint alignment-structure classification identifies 21,391 structurally supported catalog loci that a fixed human DIAMOND protein search does not confidently assign on its own, including 21,184 loci with no detectable human protein-sequence match and 207 loci with only low-confidence matches in the classical 20-30% amino-acid-identity twilight zone. These rescue-space totals describe catalog loci rather than validated one-to-one human-absent genes. In a single-cell RNA sequencing application, a stricter UTR-aware Ensembl+Evo2 reference improves gene recovery and expands the interpretable feature space of the lamprey immune compartment relative to the Ensembl baseline. This enables more resolved annotation of four transcriptionally defined immune cell states, including VLRA+-associated T-like and VLRB+-associated B-like programs together with oxidative iron-handling and iron-associated VLR-linked states. Together, these results show that structural protein signal often persists beyond the limits of pairwise sequence alignment and that an embedding-based annotation layer can extend that signal to improve downstream comparative and…
Terpenes constitute the largest and most structurally diverse class of plant secondary metabolites, with critical roles in plant-environment interactions and broad industrial applications. Although nuclear genome engineering of terpene pathways has been extensively explored, chloroplast genome engineering remains largely undeveloped, with all reported studies restricted to the model plant Nicotiana. Here we report successful chloroplast genome engineering for diterpene production in the crop plant potato (Solanum tuberosum). First, we identified the trnT/trnL plastomic locus as optimal for minimizing integration-associated growth penalties. Insertion of a bifunctional diterpene synthase gene into this plastomic site yielded transplastomic plants with successful diterpene production, but with reduced growth. The co-expression of a geranylgeranyl diphosphate synthase gene to enhance precursor supply restored normal growth while elevating diterpene accumulation. Transplastomic plants were otherwise agronomically comparable to wild-type. This work expands chloroplast engineering as a viable strategy for terpene pathway engineering in crop improvement and high-value terpene production.
Phyllanthus niruri (Phyllanthaceae) is a medicinally important herb known for producing phyllanthin, a bioactive dibenzylbutane lignan with reported hepatoprotective and antioxidant properties. However, the biosynthetic basis of phyllanthin production remains unresolved, largely due to the absence of a reference genome for the species. We report a chromosome-scale assembly of P. niruri generated by integrating PacBio HiFi long reads and Illumina short reads, followed by reference-guided scaffolding against Phyllanthus cochinchinensis. The assembly has an L50 of 7 and 97.6% BUSCO completeness. Annotation predicted 19,254 protein-coding genes, of which 91.1% were functionally annotated, with phenylpropanoid biosynthesis emerging as the most enriched specialized-metabolism pathway in the genome. Using pathway-guided genome mining, structural similarity analysis, and comparative metabolic reconstruction, we propose a putative biosynthetic pathway for phyllanthin originating from the phenylpropanoid-lignan branch through secoisolariciresinol-like intermediates followed by terminal O-methylation reactions. A total of 305 unique candidate genes associated with the proposed pathway were identified, including expanded families of dirigent proteins, peroxidases, secoisolariciresinol dehydrogenases, and O-methyltransferases. Comparative transcriptomic analyses across related Phyllanthus species further supported the proposed pathway through coordinated expression of lignan-associated genes and tissue-specific enrichment of O-methyltransferases. This work provides the first reference genome for P. niruri and a prioritized candidate gene set for functional characterization of phyllanthin biosynthesis.
Holcosus orcesi, the Orces Blue Whiptail, is a Critically Endangered lizard endemic to the upper Jubones River basin in southern Ecuador. Restricted to a narrow elevational range within semi-arid Andean shrublands, it represents one of the few montane members of a predominantly lowland lineage. Here we present the first high-quality reference genome for H. orcesi, generated using Oxford Nanopore Technologies long-read sequencing. The assembly spans 1.68 Gb across only 91 contigs, with an N50 of 76.2 Mb and a BUSCO completeness of 96.8%, making it among the most contiguous and complete squamate genomes to date. Structural annotation predicted 25,682 genes, of which 85% showed homology to known proteins and 45% were assigned Gene Ontology terms. Repetitive elements accounted for 46.3% of the genome, with LINEs representing the predominant class. This genome provides a foundational resource for future evolutionary, comparative and conservation-genomic research of H. orcesi and other mountain reptiles, enabling studies of population genomics, local adaptation, and genomic erosion in isolated populations. By expanding the genomic representation of tropical montane reptiles, this work helps address longstanding phylogenetic and geographic gaps in global biodiversity genomics and provides a foundation for evidence-based conservation of H. orcesi and related taxa.
Measles virus remains a significant global health threat, and despite the availability of an effective vaccine, measles cases continue to increase worldwide in recent years. Genomic surveillance has become an essential tool for monitoring virus circulation and investigating outbreaks. Here, we describe a wet laboratory method for whole genome sequencing of measles virus using a tiled amplicon approach and Illumina sequencing technology. A previously published Oxford Nanopore based tiled primer scheme was adapted to include both circulating measles genotypes and for use on the Illumina platform. Two Illumina library preparation kits, Illumina DNA Prep (IDP) and Nextera XT (XT), were evaluated for performance. The IDP kit demonstrated more complete genomes and consistent genome coverage compared with XT. Using quantified reference genomes, the limit of detection was determined to be 10,000 genome copies for genotype B3 and D8. Sequence accuracy was evaluated using previously characterized clinical samples and showed high concordance. This method provides a reliable and sensitive approach for measles virus whole-genome sequencing using Illumina platforms and is suitable for genomic surveillance applications.
Abstract As part of preparedness activities supporting pathogens classified under the UK High Consequence Infectious Diseases (HCID) framework, we previously evaluated both a whole-genome tiling amplicon sequencing scheme and a pan-viral hybridisation capture approach (TWIST-CVRP) for sequencing Andes virus (ANDV). In light of the recent outbreak, we make available viral sequencing datasets generated using a historical ANDV isolate (Chile, 1997). In addition, we provide an evaluation of tiling amplicon scheme performance and present recommended primer updates informed by in silico comparison with the recently released outbreak genome. These datasets are intended to support benchmarking, validation, and optimisation of bioinformatic pipelines across the community.
Biological sequences are known to be not random. Thus, the comparison of in silico restriction fragment distributions of random and biological sequences may be an indicator of this non-randomness. Our analyses show that for most of the tested combinations of restriction enzyme and genome sequence the fragments per Megabase of the biological sequence deviate at least more then 10% from the corresponding random sequence. This deviation goes into both directions, i.e. clearly increased values are as common as clearly decreased values. Although there is no species- or restriction-enzyme-specific effect, a clear impact of the GC content both of the restriction site and of the genome sequence can be seen. In contrast to the random sequences, the genome sequences show distinct peaks in their fragment length distributions, hinting to repetitive elements such as transposons.
As one of the earliest-diverging multicellular eukaryotic lineages, the bladed Bangiales (Rhodophyta) possess a deep evolutionary history with a central role in the multi-billion-dollar global seaweed aquaculture industry. Although North Atlantic representatives are emerging candidates for regional mariculture, the scarcity of high-quality genomic resources for these taxa hinders both fundamental research and commercial optimization. To address this, we present the first chromosome-level genome assemblies for two native European species: Porphyra dioica (150.44 Mbp) and Porphyra linearis (95.22 Mbp). By integrating Oxford Nanopore Technologies (ONT) long-read sequencing with Hi-C proximity ligation, we generated highly contiguous nuclear genomes resolved into five chromosomes. Structural gene models were predicted through the BRAKER3 pipeline, identifying 12,548 and 10,382 protein-coding genes for P. dioica and P. linearis, respectively. Subsequent homology-based functional annotation characterized 57.4% and 59.8% of these predicted proteins. Supplemented by circularized organellar genomes, these reference genomes provide a critical framework for future research, enabling comparative studies of Atlantic-Pacific divergence and facilitating the development of selective breeding programs for sustainable European aquaculture.
LD Score Regression (LDSC) is a prominent method, which estimates whole-genome SNP heritability from summary statistics via the slope of a linear regression of GWAS test statistics corresponding to a trait of interest against LD scores. It was claimed by the LDSC authors that the free intercept in the regression accounts for confounding bias such as population stratification. In this study, we argue that the intercept in LDSC must be fixed to 1 for accurate SNP heritability estimation. We show both theoretically and with simulations that the estimated intercept does not accurately capture population stratification effects, and that it adversely affects the accuracy of the heritability estimate introducing bias and increasing variance. Fixing the intercept to 1 eliminates bias and reduces variance when no population stratification is present. On the other hand, under population stratification, LDSC is biased with both the free and the fixed intercept. Additionally, we show that estimated standard errors in LDSC are underestimated, potentially leading to false-positives in downstream GWAS analyses.