As one of the earliest-diverging multicellular eukaryotic lineages, the bladed Bangiales (Rhodophyta) possess a deep evolutionary history with a central role in the multi-billion-dollar global seaweed aquaculture industry. Although North Atlantic representatives are emerging candidates for regional mariculture, the scarcity of high-quality genomic resources for these taxa hinders both fundamental research and commercial optimization. To address this, we present the first chromosome-level genome assemblies for two native European species: Porphyra dioica (150.44 Mbp) and Porphyra linearis (95.22 Mbp). By integrating Oxford Nanopore Technologies (ONT) long-read sequencing with Hi-C proximity ligation, we generated highly contiguous nuclear genomes resolved into five chromosomes. Structural gene models were predicted through the BRAKER3 pipeline, identifying 12,548 and 10,382 protein-coding genes for P. dioica and P. linearis, respectively. Subsequent homology-based functional annotation characterized 57.4% and 59.8% of these predicted proteins. Supplemented by circularized organellar genomes, these reference genomes provide a critical framework for future research, enabling comparative studies of Atlantic-Pacific divergence and facilitating the development of selective breeding programs for sustainable European aquaculture.
Science Journals
LINE-1 retrotransposons are the only autonomous mobile elements still active in human genomes and remain a potent source of mutation, genome remodeling, and disease risk. However, young, full-length, potentially active copies (the elements most likely to shape present-day genomes) have been largely inaccessible to population-scale analysis because they are long, repetitive, and poorly resolved by short-read sequencing. Here, we use 47 phased long-read assemblies from the Human Pangenome Reference Consortium, representing 94 haplotypes, to build an allele-resolved view of recent human LINE-1 evolution. We identify 13,617 LINE-1 alleles with intact ORF1 and ORF2 across 683 unique insertion sites, revealing that every genome carries a distinct repertoire of potentially active source elements. These intact LINE-1 profiles recapitulate broad human population structure while exposing a large, rare, and population-enriched reservoir of mobile-element diversity missed by single-reference approaches. We also resolve a structurally variable chromosome 11 LINE-1 array, demonstrating that local duplication and rearrangement can amplify LINE-1 sequence independently of canonical retrotransposition. By comparing full-length LINE-1 sequences, we define activity signatures that separate ancient remnants from recently expanding lineages and uncover young LINE-1 groups whose activity is not fully explained by canonical subfamily labels. Sequence-network analyses further reveal a dynamic history of lineage turnover, in which successful source elements rise, seed new insertions, and are replaced by descendants marked by specific nucleotide changes. Together, these data transform human LINE-1s from a repetitive background into a resolved evolutionary system, linking insertion polymorphism, coding potential, population history, and recent retrotransposon adaptation. Our findings establish the human pangenome as a framework for discovering active source elements and for testing how mob…
Elucidating how habitat degradation facilitates extinction is critical for effective conservation efforts. Here, we propose integrating physiologically-structured population models into stochastic population viability analyses to assess how differing consequences of habitat degradation interact to drive extinction dynamics in a focal population. Using the isolated spectacled caiman Caiman crocodilus population/ecomorph from the Apaporis River as a case study, we find that threatening the resource base, which individuals increasingly rely upon, to outgrow vulnerable size ranges and mature accelerates extinction. We also found that when habitat degradation impacts both the primary adult and juvenile resource bases, this can have marked synergistic effects on threatening population viability. By contrast, destroying nesting sites has only a small effect on accelerating the impact of deteriorating prey availability. Through integrating community-level feedback between habitat degradation/change and population dynamics/structure, our approach provides a comparative framework for assessing the relative importance of distinct mechanisms through which habitat degradation ultimately drives extinction risk.
Accurate and reproducible assessment of foliar disease severity is essential for evaluating the performance of heterogeneous plant communities and understanding host-pathogen interactions. However, traditional visual scoring methods remain subjective, with limited precision, and difficult to scale in large phenotyping experiments. Here, we present a semi-automated image analysis workflow designed to quantify multiple foliar disease symptoms simultaneously on wheat flag leaves sampled from varietal mixtures. The workflow combines three methodological components: (i) a standardized protocol for leaf sampling and imaging, (ii) supervised machine learning segmentation using Random Forest implemented in Ilastik to classify multiple symptoms (powdery mildew and yellow rust), and (iii) a graphical user interface facilitating pipeline deployment by non-specialist operators. To evaluate the influence of image representation on classification performance, four color spaces (RGB, HSV, HLS, LAB) were systematically compared. The approach was validated using images of durum wheat flag leaves collected from a field experiment assessing eight-way varietal mixtures under natural fungal pressure. Cross-validation against manually annotated images demonstrated high segmentation accuracy across all symptom. Comparison among color spaces revealed only minor differences in performance. Overall, this workflow offers a cost-effective, annotation-efficient and reproducible alternative to deep learning approaches, leveraging open-source and actively maintained tools while requiring limited training data and enabling objective, reproducible and scalable disease phenotyping.
Cancer progression is increasingly understood as an evolutionary process shaped not only by competition but also by cooperative interactions including those mediated through diffusible ``public goods'' (PGs). Classical evolutionary game theory predicts that PG-producing (altruistic) subclones cannot invade well-mixed populations of non-producers, creating a paradox given their observed emergence in tumors. Here, we resolve this contradiction by combining stochastic spatial simulations with an analytically tractable Moran model to study the invasion dynamics of PG-producing cells in structured populations. Starting from a single producer cell, we explicitly model stochastic PG secretion, diffusion, binding/unbinding, and cell proliferation across biologically relevant parameter ranges. We demonstrate that spatial structure fundamentally alters invasion dynamics, enabling PG producers to invade and establish even when production incurs a fitness cost. Both numerical and analytical approaches converge on a key unifying parameter, a characteristic length scale {delta}, that captures the combined effects of diffusivity, binding kinetics, and degradation. This length scale determines the spatial extent of PG availability and thus the selective advantage of producers. We identify distinct regimes: when PGs are localized (small {delta}), producers preferentially benefit and invasion is likely; when PGs are widely dispersed (large {delta}), benefits are shared and invasion approaches neutrality or is suppressed by costs. Our results highlight that invasion of cooperative traits is governed by spatially mediated resource localization rather than intrinsic fitness alone. This framework provides a mechanistic basis for understanding the emergence of cooperative subclones in tumors and suggests that modulating biophysical transport properties of signaling molecules could influence tumor evolution, metastasis, and therapeutic resistance.
Gene duplication is a major driver of evolution, yet how it generates fundamentally new molecular functions remains poorly understood. Here, we show how such novelty arose in KLMT-1, a selfish toxin that causes genetic incompatibilities in Caenorhabditis tropicalis. KLMT-1 evolved via duplication of an essential tRNA synthetase but, strikingly, lost its ancestral role in tRNA biology and translation. Instead, KLMT-1 localizes to centrosomes, where it targets Aurora kinase A (AIR-1). This innovation is mediated by a three-amino acid insertion that extends a beta-hairpin loop, enabling electrostatic interaction with a regulatory interface on the kinase. Our results demonstrate how changes in selective pressure, combined with minimal modifications in neutrally evolving regions, allow duplicated proteins to access new functional space and evolve entirely new molecular activities.
Gross chromosomal rearrangements are a hallmark of many diseases and cancers. The study of their biogenesis and the mechanisms underlying their formation is greatly facilitated by the availability of genetic reporter assays in model organisms. We present here a novel GCR assay developed in fission yeast, a highly relevant model for understanding genome instability related to human biology. The reporter employs canavanine counter-selection to detect GCRs within a chromosomal context. Using this assay, we identified natural hotspots for GCRs, including inverted long terminal repeats (IR-LTRs). Structural analysis of GCR events showed that IR-LTR-induced GCRs mainly result in either terminal deletions with adjacent inverted duplications or repair via long-range break-induced replication (BIR). Deleting IR-LTRs reduces the GCR rate and reveals another hotspot driven by BIR between homeologous aldo/keto reductase genes on opposite arms of chromosome I. This is the first evidence that BIR can occur in S. pombe on long tracks reaching up to 600 kb. Besides highlighting genome rearrangement hotspots, the assay also identifies regulators of genome instability in fission yeast. Loss of Nup132, a component of the nuclear pore complex, increases IR-LTRs-induced GCRs, while the budding yeast homolog Nup133 has no effect on the stability of a structurally similar IR. In contrast, disrupting djc9, which encodes a conserved histone H3-H4 binding protein, decreases GCR rates. Overall, this sensitive GCR assay enables the identification of factors that control spontaneous and fragile motif-induced chromosomal instability, including those conserved in humans but lost through evolution in other organisms.
Although resources are typically distributed continuously in space, species distributions often organize into discrete clusters. In his seminal paper, Turing demonstrated that such clusters can spontaneously arise in population densities, even when populations evolve in environments with continuously varying conditions. This phenomenon is known as Turing instability. In this work, we focus on two models grounded in population dynamics: a one-dimensional model based on the nonlocal Fisher-KPP equation, and a two-dimensional model involving an environmental gradient. We show that phenotypic clusters (sometimes referred to as "species") emerge in these models. We prove that they do not emerge because of Turing instability, but because of stochasticity, and that they disappear when stochasticity is reduced. First, for both models, we start our simulations with initial populations uniformly distributed in the state space. We show that phenotypic clusters quickly emerge and that the distances between them depend on the population size, that is, on the degree of stochasticity. Next, we start from already clearly defined phenotypic clusters. We identify three regimes in the connection between population size, the initial distances between clusters, and the distances between clusters at equilibrium. Last, on the two-dimensional model, we relax the hypothesis of complete clonality by varying the effective recombination rate, explore its effect on phenotypic clustering, and show that phenotypic clustering decays drastically with slight recombination.
Birds and mammals are shrinking and shapeshifting as global temperatures rise. Ecogeographic rules predict that such changes should ease heat stress by increasing surface-area-to-volume ratios, and thus, the capacity for heat exchange. This has led to the hypothesis that body size reductions are driven by thermoregulatory selection or adaptive plasticity, although recent syntheses point to more complex, multifactorial causes. Crucially, recent theoretical models predict that thermoregulatory benefits of smaller body size only emerge at extreme deviations from average phenotypes. Here, we exploit agricultural selection in Japanese quail to directly test this hypothesis, using three breeds spanning extreme differences in body mass, surface area, and relative appendage lengths. Evaporative cooling capacity and the scope for evaporative water loss broadly followed allometric predictions when contrasting small and larger breeds. As expected, this allowed the smallest breed to tolerate higher air temperatures. However, differences in heat tolerance limits between breeds were consistently much smaller than predicted. Additionally, the breadth of thermoneutral zones overlapped in full, and upper critical temperatures were remarkably similar, between breeds. Together, these results show that heat tolerance is only weakly linked to surface-area-to-volume relationships and cannot be explained by size alone. Thus, although smaller bodies may modestly enhance heat dissipation when size variation in a population is substantial, our findings suggest that recent body size reductions and morphological shifts are unlikely to be driven primarily by thermoregulatory benefits.
Anxiety has been extensively studied in relation to memory, yet its dynamic association with spatial episodic memory in naturalistic clinical settings remains largely unexplored. We developed an anxiety-spatial-memory EMA protocol (asm-EMA) and deployed it in 30 epilepsy patients undergoing inpatient EEG monitoring, delivering combined momentary anxiety ratings and a validated spatial memory task pseudo-randomly every 90-150 minutes across multiple days. Subject-level asm-EMA means and session-to-session variability both correlated significantly with standard neuropsychological assessments, supporting the clinical validity of our design. Elevated within-person STAI-6 was selectively associated with faster retrieval responses, yet spatial memory accuracy was independent of all three anxiety measures, suggesting a shift in response strategy rather than memory impairment. Within-day anxiety showed short-term carryover between consecutive sessions, with little persistence beyond the next session. The asm-EMA protocol provides a feasible, autonomous framework for capturing moment-to-moment anxiety-memory dynamics in naturalistic settings.
Wildlife vaccination could become a powerful strategy to mitigate disease-induced biodiversity losses, yet many vaccines for wildlife diseases provide only limited protection. Notably, tools to control the fungal pathogen Batrachochytrium dendrobatidis (Bd) are urgently needed for amphibian conservation. Laboratory experiments have demonstrated that prophylactic exposure to Bd metabolites increases host resistance, significantly reducing infection intensity in amphibians subsequently challenged with live Bd. Because Bd metabolites are non-infectious and applied topically, this treatment has potential to be administered to waterbodies to vaccinate and protect amphibians. We developed an agent-based model that indicated imperfect vaccination could reduce or amplify Bd infections at the population level, depending on degree of enhanced resistance or tolerance. Utilizing a Before-After-Control-Impact design with ten years of data, we conducted an ecosystem-level trial where we applied low levels of Bd metabolites or a sham control treatment to ponds in California and subsequently quantified Bd prevalence and infection intensity in metamorphosing Pacific chorus frogs (Pseudacris regilla). Unexpectedly, infection intensity was significantly greater in treated ponds relative to control ponds following metabolite addition. Additional model simulations indicated that this could occur via two mechanisms: (1) if treatment greatly increased tolerance alone or in combination with smaller increases in resistance, or (2) if a deleterious environmental interaction caused the treatment to increase susceptibility, rather than promote resistance. Future research is needed to determine whether tolerance or environmental factors drove heightened Bd infection intensities in this field trial to identify contexts in which this treatment can be used as a conservation tool.
Cliffs are environmentally extreme yet biodiversity-rich ecosystems that harbour specialist plants, many endemic and threatened. Plant persistence in these nutrient-poor substrates may depend on tightly linked soil- and root-associated microbial communities, which remain poorly understood. These interactions may become increasingly important with the global expansion of recreational climbing. While physical climbing impacts on vegetation are documented, potential chemical effects, from the use of climbing chalk (magnesium carbonate), on soil properties and plant-associated microbiota remain unknown. We sampled soils and roots beneath cliff-specialist and generalist plants, and unvegetated soils, across climbed and unclimbed routes in northern, central, and southern Spain. Soil physicochemical properties were quantified, fungal communities were characterized using ITS-metabarcoding, and structural equation modelling was used to disentangle direct and indirect effects. Climbing increased soil pH and altered soil chemical properties, driving shifts in fungal diversity and functional composition in soil and roots. The relative read abundance of root-associated symbiotrophic fungi declined, whereas arbuscular mycorrhizal fungi and pathogens increased in climbed cliffs. Overall effects were consistent, with cliff-specialist plants mediating nutrient and fungal shifts. Our findings show that climbing can reshape cliff soil chemistry and fungal communities, with potential cascading consequences for plant functional performance, nutrient dynamics, and ecosystem resilience.
Microbes that remain uncultivated occupy nearly every ecosystem on the planet; this is particularly true in soils, where despite their prevalence, the roles of rarely cultivated microbes in driving biogeochemical cycles and ecosystem function remain poorly explored. We combine metagenome-informed substrate selection with enrichment sub-communities to generate reduced-complexity communities that preserve co-occurrence and expand experimental access to underrepresented soil lineages without requiring prior isolation of each member. Carbohydrate-active enzyme (CAZyme) profiles from soil-derived genomes were used to select carbon compounds predicted to enrich difficult to culture taxa, including members of the phylum Acidobacteriota. Based on 16S rRNA amplicon sequencing, we reproducibly enriched Terriglobus (Acidobacteriota) on multiple metagenome-guided substrates. Select communities with consistent presence and varying abundance of Terriglobus were passaged in a longitudinal design to generate 89 metagenomes; genus-level profiling revealed that community composition varied between biological replicates but remained consistent within replicates over time, providing diverse Acidobacteriota-containing configurations for downstream analysis. Association network inference identified a core set of co-occurring taxa that positively tracked with Terriglobus across the longitudinal series. In parallel, the substrate-guided approach led to isolation of a novel Terriglobus species, the first cultured representative of its GTDB species cluster. Together, these results establish a generalizable strategy for generating communities enriched with rarely cultivated taxa, yielding tractable systems for studying microbial interactions and community assembly in soil.
Ocean warming is altering abiotic environments and biotic interactions experienced by marine organisms, where sensitive early developmental windows occur in biologically complex seawater communities. The impact of these interactions on developmental processes and fitness in hosts is not well understood, but likely contingent on the establishment of a host-associated microbiome. Here, we hypothesize that temperature and microbial exposure during embryogenesis influence larval microbiome assembly and host morphology. Strongylocentrotus purpuratus embryos were raised in low microbial richness (LMR) or high microbial richness (HMR) seawater at ambient (14 {ring}C) or elevated (18 {ring}C) temperature, then collected at 2, 4, and 6 days post-fertilization (dpf) following multiple feedings. Higher microbial diversity was observed in larvae that developed in HMR seawater when compared to LMR. Differences in relative abundances of dominant microbial families between seawater and larvae suggest some degree of host selectivity in microbiome assembly. Temperature did not strongly alter microbiome composition, but both temperature and microbial condition led to differences in larval morphology by 6 dpf, potentially due to enrichment of microbes with chemoheterotrophic functions. By linking how temperature and microbial communities interact with host development, we contribute novel insights into how early-life environmental conditions impact holobiont formation and morphology.
Mapping the genetic basis of inter-individual heterogeneity in multifactorial diseases opens the door to mechanistic insights and opportunities for targeted intervention. In Alzheimer's disease (AD), clinical and pathological heterogeneity is well recognized, but genetic dissection is limited by a lack of well-powered cohorts with deep phenotypic characterization. Here, we introduce a polygenic score (PGS) analysis strategy to address these limitations by leveraging the inherent pleiotropy in complex trait genetics. We perform a cross-cohort, cross-trait application of pre-trained PGS, integrating 713 UK Biobank-derived PGS with 36 deep AD phenotypes across 1678 ROSMAP participants. We identify 268 statistically significant (FDR<0.1) associations between 12 prioritized PGS and 36 AD phenotypes. Prioritized PGS include blood lipid measurements, inflammatory biomarkers, and cancer traits; observed AD phenotypes include cognition, amyloid, and tangles. Of the 268 associations, 49 persist with APOE-excluded PGS. Predictive models trained on multiple prioritized PGS outperform the AD PGS or APOE alone for predicting amyloid and cognition. Lastly, our approach identifies six individual-level AD polygenic subtypes supported by distinct pathological patterns. Overall, we combine large-scale biobank resources and deeply-phenotyped cohorts using PGS, reveal genetic features underlying AD heterogeneity, and provide a general model for stratifying heterogeneous disease-focused cohorts using genomics.
Despite sharing the same genes and the same environment, individuals often develop substantial phenotypic differences. While this pattern has been documented across diverse species and traits, the processes giving rise to this 'stochastic' or non-shared environmental variation remain unclear. Recent mathematical models of development in which phenotypes are gradually constructed may offer some clues. These models show that imperfect environmental cues can generate striking variation in developmental trajectories and adult phenotypes. At the population level, such imperfect cues produce increasing stability of individual differences across ontogeny (e.g. animal personality) and patterned distributions of mature phenotypes (e.g. normal or skewed) that resemble those observed in real organisms. Our paper synthesizes existing models in which stochastic phenotypic variation arises solely as a by-product of mechanisms missing their phenotypic targets because of imperfect cues. We then link these models to related, but independent, mathematical theory exploring the environmental conditions under which stochastic phenotypic variation is favoured by natural selection. Our integration shows that stochastic sampling is often favoured over classic bet-hedging strategies involving non-plastic generalist or specialist strategies. Our findings provide new directions of research on stochastic sampling as a mechanism for adaptive stochastic variation within and across generations.
How do humans store sequences that far exceed working memory capacity? Using visuo-spatial and binary auditory sequences, we previously showed that a Language of Thought (LoT) architecture, in which simple primitives are recursively combined into hierarchical programs, enables efficient storage of structured sequences. Here we ask whether this principle extends to purely ordinal structure: sequences defined by how items repeat and in what order, as in AABBCCAABBCC, independently of their spatial content. Across three experiments, participants reproduced 12-item sequences of spatial locations with various ordinal structures. The minimal description length derived from the LoT model predicted recall accuracy with remarkable precision (r = .96), substantially outperforming Shannon entropy, Lempel-Ziv complexity, chunking models and subjective complexity ratings. Critically, fine-grained analyses of participants' inter-click intervals during reproduction revealed systematic slowdowns at the hierarchical boundaries predicted by the LoT programs, providing a behavioral signature of the underlying mental syntax. These results identify a compact vocabulary of mental primitives, repetition, mirroring, and interleaving, whose composition accounts for the symbolic compression of ordinal structures. For ordinal regularities, human sequence memory operates as a form of program induction, leveraging a domain-general capacity for hierarchical compression to encode complex structured information.
Humans comprehend language incrementally, updating the representation of sentence meaning with each incoming word. These updates are guided by the distance between each perceived word and prior expectations--the prediction error. The alignment between large language models (LLMs) and cortical activity inspires the hypothesis that the cortical computation of prediction error is Surface-based, driven by statistical patterns of word form co-occurrence. In contrast, psycholinguistic models propose that prediction error computation is Meaning-based, driven by word semantics. We used polysemic words with ambiguous semantics to distinguish these models: ambiguity would introduce uncertainty into meaning representations and hence the prediction error, if Meaning-based, but would not affect the prediction error, if Surface-based. We examined how ambiguity influenced prediction error signatures in self-paced reading times and magnetoencephalographic (MEG) neural responses during sentence processing. While an LLM-based proxy of prediction error robustly predicted reading times and neural responses to unambiguous words, it failed to predict either under ambiguity. That is, prediction error computation was altered by uncertainty in word meaning, which supports the Meaning-based model and corroborates the essential role of word meaning in predictive language processing. Our findings highlight an important limitation of LLMs as in silico models of the human language faculty.
Metatranscriptomic (MTX) sequencing enables profiling of gene expression across microbial communities, providing a framework for linking genetic potential with functional activity. However, standard pipelines report normalized abundances rather than raw counts, limiting the use of count-based RNA-seq methods, while Gaussian-based alternatives rely on transformations and assumptions that are often poorly suited to MTX data. We propose a new modeling framework for differential expression analysis of MTX data, built on a scale mixture of exponential distributions, that incorporates DNA abundance to adjust for genomic potential, accommodates subject-specific random effects, treats zeros as left-censored, and employs a mixture prior to handle extreme sparsity. Applied to the IBDMDB multi-omics cohort, differential expression results vary substantially across models, including among Gaussian approaches with different pseudocount choices. Our approach identifies a distinct subset of candidate genes not detected by existing Gaussian methods; these may provide useful leads toward a novel understanding of transcriptomic patterns associated with dysbiosis in inflammatory bowel disease. Estimated dysbiosis effect directions are consistent between our model and Gaussian-based approaches, while effect sizes from our model tend to be larger in absolute value.
Transcriptomics has transformed our understanding of the brain, but assigning transcriptomic identities to neurons recorded in vivo remains challenging at scale. Existing platforms can pair transcriptomic identity with two-photon calcium imaging in small populations of approximately 100 neurons, but they require recorded cells to be sparse and therefore cannot be applied to large population recordings. Here, we present coppaFISH 3D, a spatially resolved transcriptomics method, and CASTalign, an in silico alignment framework, which together enable transcriptomic identification of thousands of simultaneously recorded cells. coppaFISH 3D detects hundreds of genes in thick 50m fixed sections while preserving tissue integrity, enabling both 3D registration to in vivo imaging and integration with immunofluorescence labelling. The platform is fully powered by open chemistry and open source software, runs on commodity hardware, and can be performed at very low cost per section. It therefore enables transcriptomic identification of recorded neurons at scale, making it possible to study how transcriptomic identity shapes activity in neural populations.
How post-mitotic neurons maintain precise transcription factor (TF) levels throughout life remains a fundamental open question. Here, we challenge the prevailing model of positive autoregulation by demonstrating that UNC-3 (Collier/EBF1-4), a dosage-sensitive TF continuously required for cholinergic motor neuron identity in C. elegans, negatively regulates its own expression. Using genetics, biochemistry, and inducible protein depletion, we show this self-repression occurs directly at the transcriptional level and persists beyond development. CRISPR/Cas9 disruption of negative autoregulation causes motor neuron identity and locomotion defects, establishing its functional necessity. Mechanistically, the UNC-3 DNA-binding domain is required and sufficient for self-repression, with an AlphaFold2 screen implicating chromatin factors as interaction partners. Critically, UNC-3 self-repression is continuously counterbalanced by positive input from the HOX cofactor CEH-20/PBX, revealing a dynamic "balancing act" between opposing regulatory inputs that stabilize TF dosage over time. Mutations in the unc-3 ortholog EBF3 cause a neurodevelopmental syndrome, and disease-associated variants disrupt UNC-3 self-repression, revealing a key molecular mechanism underlying the disorder. We propose that negative autoregulation continuously counteracted by positive input represents a broadly applicable principle for maintaining dosage-sensitive TF expression to secure post-mitotic cell identity.
Conformational plasticity of RNAs plays important roles in recognizing RNA-binding proteins, and is often modulated by their binding partners. Here, we investigate RNA conformational preferences in a non-redundant dataset of 263 protein-RNA complexes to characterize the structural landscape associated with protein recognition. RNA dinucleotide segments are analyzed using seven backbone torsion angles ({delta}1, {varepsilon}1, {zeta}1, 2, {beta}2, {gamma}2, and {delta}2), two glycosidic torsion angles ({chi}1 and {chi}2) and the pseudo-torsion angle . Focusing on dinucleotide steps present in both interface and non-interface regions, we performed density-based clustering using selected backbone torsion angles to identify recurrent conformational states. We identify 28 distinct RNA dinucleotide conformers containing at least ten members each. Among these, eight conformers represent previously unreported nucleotide conformers (NtCs), including the transitional and the non-canonical states AB06, AB07, BB21, BB22, OP32, OP33, IC08 and IC09. Several of these conformers are preferentially enriched at protein-binding interfaces, suggesting their involvement in local conformational adaptation during protein-RNA recognition. The newly identified conformers span transitional A-B geometries, distorted B-like states, open conformations and compact intercalated structures, highlighting the remarkable structural plasticity of RNA in ribonucleoprotein complexes. Overall, this study expands the current understanding of RNA conformational space and provides a refined RNA dinucleotide conformer library for protein-RNA complexes. These findings will facilitate the identification of novel RNA structural motifs and improved RNA structural modeling, docking protein-RNA complexes and deep learning-based prediction frameworks for describing RNA tertiary structures.
Coxiella burnetii is the only member of the order Legionellales known to primarily infect vertebrates. The Q fever pathogen is also unusual in that it replicates within an acidified phagolysosome-like vacuole. The evolutionary origins of the virulence determinants underlying this lifestyle remain unclear. More broadly, little is known about how virulence-related traits arise in specialized intracellular lineages, where access to foreign-origin DNA may be more episodic. To address this question, we used Legionellales-wide comparative phylogenomics to reconstruct the gain and loss of traits affecting host interaction, immune evasion, intracellular survival, and metabolism. We found that many virulence-associated traits in C. burnetii predate the modern pathogen and were assembled stepwise in ancestors that likely occupied niches distinct from the acidified vacuolar niche of modern C. burnetii. The common ancestor shared with soft-tick Coxiella endosymbionts likely encoded most C. burnetii type IVB secretion system effectors, indicating that much of the host-manipulation repertoire in C. burnetii was already present before the emergence of the modern pathogen. Distinctive lipopolysaccharide features associated with immune evasion also appear to have accumulated progressively within the Coxiella lineage, including genes implicated in synthesis of virenose, a unique O-antigen sugar critical for C. burnetii virulence. Traits likely to support replication in the acidic Coxiella-containing vacuole likewise accumulated gradually, with generalized stress-tolerance functions predating acquisition of an Mrp cation/proton antiporter that may have further supported pH homeostasis. Additional changes in sugar transport and catabolism, glycolytic control, and respiratory metabolism may have enhanced metabolic flexibility and access to diverse substrates in this nutrient-rich niche. Together, these findings support a model in which vertebrate pathogenicity in C. burnetii emerged …
Short-read amplicon sequencing is widely used for fungal surveys but can limit taxonomic resolution. Long-read sequencing enables recovery of the full internal transcribed spacer (ITS) region and may improve ecological and taxonomic inference. Here, we conducted a paired comparison of Illumina ITS2 and PacBio HiFi full-length ITS sequencing using identical DNA extracts from built-environmental air and surface samples (n = 68) collected across homes, a dormitory, and laboratories. Both datasets were taxonomically assigned using the same algorithm and reference database. We performed paired statistics, in-silico ITS2 trimming of long-read sequences, and cross-platform mapping at multiple identity thresholds. Full-length ITS provided higher taxonomic resolution, assigning a greater fraction of ASVs at the family (98% vs. 88%) and species (42% vs. 32%) ranks than ITS2 (paired Wilcoxon q=0.002). Alpha-diversity comparisons showed similar Shannon diversity across pipelines, whereas richness metrics were consistently higher for full-length ITS. Beta-diversity analyses indicated broadly comparable community-level patterns, although full-length ITS revealed stronger sample-type- and location-associated structure (PERMANOVA R{superscript 2} 0.06, p=0.0001). In-silico ITS2 trimming reduced these differences, indicating that amplicon length is a major contributor to enhanced taxonomic resolution and ecological inference. Cross-platform mapping further showed extensive one-to-many relationships between ITS2 and full-length ITS ASVs, consistent with increased sequence resolution in long-read data.Together, these results show that ITS2 sequencing provides robust community-level profiling, while full-length ITS enables improved richness estimates and finer ecological and taxonomic resolution. This paired, bias-aware framework provides a practical template for selecting fungal amplicon sequencing strategies in built-environment mycobiome studies.
Course-based undergraduate research experiences (CUREs) can expand undergraduates' access to research and motivate students to stay in science. Yet, little research has examined how CURE instruction shapes student motivation. We leveraged a motivation-related characterization of non-content talk of 48 CURE and non-CURE instructors to predict the motivation-related outcomes of 462 students. We fit a series of multi-level models (MLM) in which we regressed students' post-course scientific self-efficacy, task values, scientific identity, and science-related intentions onto instructors' self-efficacy and task values-related talk, controlling for students' pre-course levels. We also fit an MLM to explore whether instructors' relationship-building talk (immediacy talk) was associated with students' rapport with their instructor. Instructors' self-efficacy talk did not affect students' self-efficacy, and instructors' immediacy talk had a marginally positive but non-significant association with students' rapport ratings. Instructors' task values talk positively influenced students' scientific identity and some but not all of their task values. Instructors' task values talk also positively influenced students' intentions to pursue a science career, but not graduate education or research careers. Collectively, these results suggest that instructors' task values talk may underpin some of the motivational effects of CURE instruction, but that task values talk need not be limited to CUREs.