Biological sequences are known to be not random. Thus, the comparison of in silico restriction fragment distributions of random and biological sequences may be an indicator of this non-randomness. Our analyses show that for most of the tested combinations of restriction enzyme and genome sequence the fragments per Megabase of the biological sequence deviate at least more then 10% from the corresponding random sequence. This deviation goes into both directions, i.e. clearly increased values are as common as clearly decreased values. Although there is no species- or restriction-enzyme-specific effect, a clear impact of the GC content both of the restriction site and of the genome sequence can be seen. In contrast to the random sequences, the genome sequences show distinct peaks in their fragment length distributions, hinting to repetitive elements such as transposons.
Science Journals
Avoiding fertilization with genetically incompatible partners, whether too similar or too divergent, is a central challenge for sexually reproducing organisms. Selection can favor mechanisms acting before and after mating, with postmating processes potentially compensating for constraints on premating choice. In the postmating context, female reproductive fluid (FRF) can modulate sperm performance and bias fertilization outcomes, but its contribution to reproductive isolation remains unclear. We tested whether FRF mediates discrimination against heterospecific and related sperm in two naturally hybridizing sister species of swordtails, Xiphophorus birchmanni and X. malinche, that diverge in premating behavior towards heterospecifics. Effects of FRF differed sharply between species. In X. malinche, FRF enhanced the velocity of conspecific sperm relative to heterospecifics, consistent with postmating discrimination against hybridization. In contrast, FRF in X. birchmanni did not favor conspecific sperm. Evidence for inbreeding avoidance was weaker, and we found no indication of a trade-off between discrimination against genetically similar and dissimilar sperm. These results show that female reproductive fluid can serve as a rapidly evolving axis of reproductive isolation through postmating female choice.
Reef restoration practitioners aim to preserve coral genetic diversity by protecting reefs and cultivating diverse genotypes in coral nurseries. However, cryptic genetic lineages in most corals complicate restoration strategies, as the role of between-lineage genetic divergence remains unclear regarding adaptation. In Montastraea cavernosa, researchers have identified cryptic lineages, some strongly segregated by depth. We conducted a ten-week reciprocal transplantation experiment using two cryptic lineages restricted to shallow water (<10m depth), with one lineage more common on nearshore reefs and the other on offshore reefs. We aimed to quantify lineage-specific responses to the environment that explain the genetic and ecological divergence between the two lineages. Surprisingly, the strongest response was not lineage-specific. Instead, both lineages exhibited strong and similar changes in growth and metabolomic profiles, depending on the transplantation habitat. These results suggest that cryptic lineages employ similar mechanisms of adaptation and acclimatization to environmental challenges, despite their genetic distinction.
Nature-Based Solutions are increasingly promoted to address current urban challenges. While their potential effects on vector-borne disease risks have been documented, data on Aedes albopictus, a known arbovirus vector, remain limited in France. A previous study showed that urban vegetation moderately increases the abundance of adult mosquitoes of this species, but the monitoring period lasted only six months. Using ovitraps, we monitored Ae. albopictus egg density dynamics over multiple years (2022 to 2024) and analysed its environmental predictors in various urban environments. We included lagged meteorological variables, land cover metrics, and the cumulated egg densities recorded in the previous weeks as environmental predictors. Both parametric (GLMM) and non-parametric (Random Forest) models were fitted to weekly egg counts per trap. Our findings highlight that (i) egg density dynamics were related to how vegetation classes structured the landscape, (ii) growing degree days and cumulated number of eggs recorded in specific lagged time windows were the main contributors to egg density, and (iii) the non-parametric and parametric models performed similarly in terms of prediction accuracy.
Susceptibility to viral infection varies widely but is not fully explained by genetics, immune status, or exposure level. We show that time of day strongly influences infection outcome, with up to 100-fold differences in enteric viral burden depending on infection timing. This temporal gating is abolished in mice lacking a functional circadian clock. We identify the antiviral transcription factor IRF1 as a direct target of the circadian transcription factor BMAL1, resulting in rhythmic expression of a basal antiviral gene program prior to infection. Loss of IRF1 eliminates this program and abrogates time-of-day dependent differences in viral replication. This circuit operates within intestinal myeloid cells, establishing a preexisting antiviral state. These findings indicate that the circadian clock programs host susceptibility in the intestine, before infection occurs.
Allostery enables proteins to couple environmental signals to functional outputs, yet how allosteric mechanisms diversify during evolution remains poorly understood. Here, we address this question in the ubiquitous and functionally diverse arsenic repressor (ArsR) superfamily by integrating information-theoretic bioinformatics, structural characterization of DNA recognition and NMR measurements of fast internal dynamics. We identify conserved residues that define the structural scaffold of ArsR proteins and subfamily-specific positions that encode inducer and DNA specificity. In the persulfide sensor SqrR, the crystal structure of the DNA-bound complex reveals how operator specificity is encoded by a limited set of residues, consistent with sequence-derived predictions functionally validated by in vitro transcription assays across divergent ArsR regulators. We further show that allosteric inhibition of DNA binding in SqrR occurs without large-scale conformational rearrangements and is instead associated with changes in internal dynamics, as previously observed for the zinc sensor CzrA. Together, these results support a model in which conformational entropy preserves allosteric connectivity while relaxing sequence constraints, thereby enabling functional diversification within a protein superfamily.
Cliffs are environmentally extreme yet biodiversity-rich ecosystems that harbour specialist plants, many endemic and threatened. Plant persistence in these nutrient-poor substrates may depend on tightly linked soil- and root-associated microbial communities, which remain poorly understood. These interactions may become increasingly important with the global expansion of recreational climbing. While physical climbing impacts on vegetation are documented, potential chemical effects, from the use of climbing chalk (magnesium carbonate), on soil properties and plant-associated microbiota remain unknown. We sampled soils and roots beneath cliff-specialist and generalist plants, and unvegetated soils, across climbed and unclimbed routes in northern, central, and southern Spain. Soil physicochemical properties were quantified, fungal communities were characterized using ITS-metabarcoding, and structural equation modelling was used to disentangle direct and indirect effects. Climbing increased soil pH and altered soil chemical properties, driving shifts in fungal diversity and functional composition in soil and roots. The relative read abundance of root-associated symbiotrophic fungi declined, whereas arbuscular mycorrhizal fungi and pathogens increased in climbed cliffs. Overall effects were consistent, with cliff-specialist plants mediating nutrient and fungal shifts. Our findings show that climbing can reshape cliff soil chemistry and fungal communities, with potential cascading consequences for plant functional performance, nutrient dynamics, and ecosystem resilience.
Wildlife vaccination could become a powerful strategy to mitigate disease-induced biodiversity losses, yet many vaccines for wildlife diseases provide only limited protection. Notably, tools to control the fungal pathogen Batrachochytrium dendrobatidis (Bd) are urgently needed for amphibian conservation. Laboratory experiments have demonstrated that prophylactic exposure to Bd metabolites increases host resistance, significantly reducing infection intensity in amphibians subsequently challenged with live Bd. Because Bd metabolites are non-infectious and applied topically, this treatment has potential to be administered to waterbodies to vaccinate and protect amphibians. We developed an agent-based model that indicated imperfect vaccination could reduce or amplify Bd infections at the population level, depending on degree of enhanced resistance or tolerance. Utilizing a Before-After-Control-Impact design with ten years of data, we conducted an ecosystem-level trial where we applied low levels of Bd metabolites or a sham control treatment to ponds in California and subsequently quantified Bd prevalence and infection intensity in metamorphosing Pacific chorus frogs (Pseudacris regilla). Unexpectedly, infection intensity was significantly greater in treated ponds relative to control ponds following metabolite addition. Additional model simulations indicated that this could occur via two mechanisms: (1) if treatment greatly increased tolerance alone or in combination with smaller increases in resistance, or (2) if a deleterious environmental interaction caused the treatment to increase susceptibility, rather than promote resistance. Future research is needed to determine whether tolerance or environmental factors drove heightened Bd infection intensities in this field trial to identify contexts in which this treatment can be used as a conservation tool.
Elucidating how habitat degradation facilitates extinction is critical for effective conservation efforts. Here, we propose integrating physiologically-structured population models into stochastic population viability analyses to assess how differing consequences of habitat degradation interact to drive extinction dynamics in a focal population. Using the isolated spectacled caiman Caiman crocodilus population/ecomorph from the Apaporis River as a case study, we find that threatening the resource base, which individuals increasingly rely upon, to outgrow vulnerable size ranges and mature accelerates extinction. We also found that when habitat degradation impacts both the primary adult and juvenile resource bases, this can have marked synergistic effects on threatening population viability. By contrast, destroying nesting sites has only a small effect on accelerating the impact of deteriorating prey availability. Through integrating community-level feedback between habitat degradation/change and population dynamics/structure, our approach provides a comparative framework for assessing the relative importance of distinct mechanisms through which habitat degradation ultimately drives extinction risk.
Although resources are typically distributed continuously in space, species distributions often organize into discrete clusters. In his seminal paper, Turing demonstrated that such clusters can spontaneously arise in population densities, even when populations evolve in environments with continuously varying conditions. This phenomenon is known as Turing instability. In this work, we focus on two models grounded in population dynamics: a one-dimensional model based on the nonlocal Fisher-KPP equation, and a two-dimensional model involving an environmental gradient. We show that phenotypic clusters (sometimes referred to as "species") emerge in these models. We prove that they do not emerge because of Turing instability, but because of stochasticity, and that they disappear when stochasticity is reduced. First, for both models, we start our simulations with initial populations uniformly distributed in the state space. We show that phenotypic clusters quickly emerge and that the distances between them depend on the population size, that is, on the degree of stochasticity. Next, we start from already clearly defined phenotypic clusters. We identify three regimes in the connection between population size, the initial distances between clusters, and the distances between clusters at equilibrium. Last, on the two-dimensional model, we relax the hypothesis of complete clonality by varying the effective recombination rate, explore its effect on phenotypic clustering, and show that phenotypic clustering decays drastically with slight recombination.
Biodiversity is commonly summarized by macroecological mean patterns, most prominently the species-area relationship (SAR) linking habitat area to expected species richness. Yet conservation, policy, and economic decisions increasingly require risk metrics: probabilities of rare but consequential biodiversity shortfalls, including local collapse. Such tail risks are central in finance and insurance but remain difficult to quantify in ecology because the data needed to estimate full richness distributions are rarely available at decision scales. Here we provide a mechanistic route from species-area relationships to biodiversity risk metrics. We show that when regional species abundances are well approximated by Fisher's log-series, a minimal immigration-extinction mechanism yields a closed-form stationary distribution for local richness whose structure tightly couples the mean SAR to richness variability and lower-tail probabilities. This coupling implies exact fluctuation-response identities and an explicit integral transform that reconstructs collapse probabilities and other tail risk measures directly from the mean SAR. These results define ecological analogues of financial risk metrics---such as collapse probability and lower-tail quantiles---without requiring direct estimation of the full richness distribution. Using high-resolution ForestGEO tree censuses spanning tropical, subtropical, and temperate forests, we find empirical support for these predictions across spatial scales. Together, our results show how widely measurable species-area relationships can be elevated from descriptive averages to operational tools for biodiversity risk assessment and reliability-based conservation planning.
The evolution of reproductive isolation is central to speciation, yet the earliest stages of this process remain poorly understood. In particular, it is unclear how rapidly barriers to mating arise during adaptation, whether they accumulate predictably, and how they depend on ecological context. Here, we investigate the evolution of mating efficiency during prolonged asexual adaptation in diploid Saccharomyces cerevisiae. Twelve replicate populations were evolved for 1200 generations in two distinct carbon environments, glucose and galactose, under strictly asexual conditions. At regular intervals, we induced sporulation and quantified mating efficiency using three complementary assays: within-population crosses, crosses between populations evolved in different environments, and crosses between evolved populations and the ancestral strain. We find that mating efficiency evolves during asexual adaptation, with outcomes that depend strongly on the environment. While glucose-evolved populations remain largely stable, galactose-evolved populations exhibit a reversible decline. Overall, changes in mating efficiency are dynamic, heterogeneous, and often transient, with evidence for both intrinsic reductions in mating competence and context-dependent incompatibilities between populations. Together, these results show that asexual adaptation can generate rapid but non-monotonic changes in mating compatibility. Early reductions in mating efficiency are heterogeneous, environment-dependent, and often reversible, and do not accumulate into stable reproductive isolation over the timescale examined. Our findings suggest that the initial stages of divergence are characterized by dynamic and contingent perturbations of reproductive traits, rather than a steady progression toward speciation.
Cancer progression is increasingly understood as an evolutionary process shaped not only by competition but also by cooperative interactions including those mediated through diffusible ``public goods'' (PGs). Classical evolutionary game theory predicts that PG-producing (altruistic) subclones cannot invade well-mixed populations of non-producers, creating a paradox given their observed emergence in tumors. Here, we resolve this contradiction by combining stochastic spatial simulations with an analytically tractable Moran model to study the invasion dynamics of PG-producing cells in structured populations. Starting from a single producer cell, we explicitly model stochastic PG secretion, diffusion, binding/unbinding, and cell proliferation across biologically relevant parameter ranges. We demonstrate that spatial structure fundamentally alters invasion dynamics, enabling PG producers to invade and establish even when production incurs a fitness cost. Both numerical and analytical approaches converge on a key unifying parameter, a characteristic length scale {delta}, that captures the combined effects of diffusivity, binding kinetics, and degradation. This length scale determines the spatial extent of PG availability and thus the selective advantage of producers. We identify distinct regimes: when PGs are localized (small {delta}), producers preferentially benefit and invasion is likely; when PGs are widely dispersed (large {delta}), benefits are shared and invasion approaches neutrality or is suppressed by costs. Our results highlight that invasion of cooperative traits is governed by spatially mediated resource localization rather than intrinsic fitness alone. This framework provides a mechanistic basis for understanding the emergence of cooperative subclones in tumors and suggests that modulating biophysical transport properties of signaling molecules could influence tumor evolution, metastasis, and therapeutic resistance.
Most species are geographically structured, leaving characteristic signatures in neutral regions of the genome. These signatures can be distorted when neutral regions are linked to deleterious mutations. In such regions, purifying selection can reduce genetic diversity through Background Selection (BGS) or, for recessive mutations, increase diversity through Associative Overdominance (AOD). While the effect of BGS and AOD are well characterized in panmictic populations, their effects remain largely unexplored in structured populations. Here, we investigated an Isolation with Migration model using forward simulations across a range of migration, selection, dominance, and recombination parameters. We first used a genotype-based approach to quantify the effects of deleterious mutations on standard summary statistics ({pi}, Dxy, FST, DAFi). We then showed that an Ancestral Recombination Graph-based approach, tracking tree sequences from a sample of one diploid per deme, recovers the same patterns while directly relating genetic variation to the underlying coalescent processes. When recombination is sufficiently low, we found a BGS-driven regime for weakly co-dominant mutations, characterized by lower diversity and increased genetic differentiation (FST). For recessive mutations, we first identified an AOD-driven regime, characterized by increased diversity and lower FST values followed by a transition to a subsequent BGS-driven regime. Genealogies were similarly impacted by deleterious mutations: BGS shrunk coalescent times and produces a shift towards lineage sorting topologies, while AOD stretched coalescent times and produces a shift toward incomplete lineage-sorting topologies. These patterns were weakened by gene flow, with FST and topologies remaining close to expected under neutrality, while diversity and coalescence times remained robust to demography. Our results provide clear evidence of BGS, AOD, and of their transition in a structured model with gene flow…
Linkage disequilibrium (LD) makes causal GWAS variants indistinguishable from correlated neighbours; resolving them is the fine-mapping problem, and the challenge is species-specific: humans face dense ancestry-imbalanced LD, yeast and *Arabidopsis* exceptionally long LD, and crop germplasm sparse and fragmented annotations that defeat human-biobank curation pipelines. Bayesian fine-mappers integrate annotations as flat per-variant priors, discarding the relational structure linking variants to tissue-specific eQTLs, pathways and protein-protein interactions. Hierarchical belief propagation (HBP) on a variant-gene-pathway factor graph matches Bayesian baselines at 5-40x speed; an annotation-adaptive complement, graph-augmented fine-mapping (GAFM), wins 27-2 against SuSiE at weak signal and recovers *LDLR*, *APOE*, *LPL*, *GCKR* and *ANGPTL3* at single-variant resolution across four Pan-UK Biobank ancestries. On the 3,000 Rice Genomes grain weight + shape panel, mixture-prior posterior reweightings of GAFM/HBP and their ensemble (GAFM-MX, HBP-MX, ENS) reach 47.6% top-1-PIP exact-position recovery of 21 panel-matched stable QTNs - the highest of any method, exceeding SuSiE (28.6%) and SBayesRC (14.3%) - at 200-700x SuSiE's per-locus speed. Across 692 leads in four species, a non-uniform per-variant prior, not uniform high coverage, lets the graph break LD ties: adding a regulatory-element flag to an otherwise uniform human cache flips HBP narrower than GAFM from 0% to 88% on 321 Pan-UKB leads. These results recast multi-omics fine-mapping as a non-uniform-prior-curation problem rather than a uniform-coverage problem, and reframe post-GWAS analysis as message passing over biological structure rather than weighted regression on flattened annotations.
Rapid connectivity alterations of thalamic nuclei during initial learning of goal-directed behaviour
The thalamus is essential for learning, dynamically engaging with other subcortical and cerebral cortex regions throughout the learning process. Here, the thalamus serves as a critical connector hub and synchroniser within the thalamocortical system of the brain. However, whilst higher order thalamic nuclei are known to be particularly important for this process, the exact contributions of individual higher order and first order thalamic nuclei, alongside their individual involvement with cortical networks and subcortical regions, remains unexplored within the initial phase of learning. In light of this, we analysed fMRI data obtained within a paradigm which is designed to examine initial learning processes within feedback-driven stimulus-response learning, in order to explore thalamic contributions. We investigated dynamic learning-related functional connectivity alterations between various thalamic nuclei with other subcortical regions and cortical networks. Our results show that the initial phase of learning was associated with: (1) decreasing functional connectivity between thalamic nuclei and frontoparietal and cingulo-opercular networks, (2) increasing functional connectivity between thalamic nuclei with default mode and salience networks, (3) decreasing functional connectivity between thalamic nuclei and the putamen, and (4) decreasing functional connectivity amongst higher order thalamic nuclei. Furthermore (5) these dynamic alterations were associated primarily by mediodorsal thalamus. Altogether, these results indicate that higher order thalamic nuclei play a crucial role within initial learning and in the generation of novel goal-directed behaviour. This was demonstrated through enhanced functional connectivity with selected cortical networks which drive goal-directed behaviour, alongside decreased functional connectivity with striatal regions which drive motor selectivity.
The hippocampus is organized along dorsal-ventral and left-right axes, but whether and how these axes interact within defined neuronal populations across behavioral states remains unresolved. Here, we combined within-animal slice electrophysiology with dual-site fiber photometry to compare dorsal and ventral CA1 activity across contralateral hemispheric configurations in mice expressing CaMKII-jGCaMP8s and SynI-jRCaMP1b at distinct longitudinal sites. Ventral CA1 pyramidal neurons exhibited greater intrinsic excitability and stronger AMPAR-mediated synaptic responses than dorsal CA1 neurons. In vivo, CaMKII-defined pyramidal recordings during home cage rest revealed a left-biased event-rate asymmetry within dorsal but not ventral CA1, with no comparable asymmetry in pan-neuronal SynI recordings. Apparent dorsal-ventral differences in spontaneous event rate were therefore configuration-dependent and resolved into a hemispheric, cell-type-specific effect restricted to the CaMKII-defined population. Lead-lag analysis showed that dorsal-ventral temporal coordination was likewise reorganized across configurations and was restricted to pyramidal-cell-biased recordings. During open-field center entries, dorsal CA1 was preferentially recruited before entry across both configurations, whereas non-coordinated entries revealed a relative post-entry suppression of contralateral ventral CA1. Together, these findings suggest that dorsal-ventral CA1 organization cannot be inferred from hemisphere-pooled designs and identify a pyramidal-cell-specific left dorsal CA1 asymmetry as a structural feature that shapes both spontaneous activity and behaviorally driven recruitment along the longitudinal hippocampal axis.
Humans comprehend language incrementally, updating the representation of sentence meaning with each incoming word. These updates are guided by the distance between each perceived word and prior expectations--the prediction error. The alignment between large language models (LLMs) and cortical activity inspires the hypothesis that the cortical computation of prediction error is Surface-based, driven by statistical patterns of word form co-occurrence. In contrast, psycholinguistic models propose that prediction error computation is Meaning-based, driven by word semantics. We used polysemic words with ambiguous semantics to distinguish these models: ambiguity would introduce uncertainty into meaning representations and hence the prediction error, if Meaning-based, but would not affect the prediction error, if Surface-based. We examined how ambiguity influenced prediction error signatures in self-paced reading times and magnetoencephalographic (MEG) neural responses during sentence processing. While an LLM-based proxy of prediction error robustly predicted reading times and neural responses to unambiguous words, it failed to predict either under ambiguity. That is, prediction error computation was altered by uncertainty in word meaning, which supports the Meaning-based model and corroborates the essential role of word meaning in predictive language processing. Our findings highlight an important limitation of LLMs as in silico models of the human language faculty.
Conformational plasticity of RNAs plays important roles in recognizing RNA-binding proteins, and is often modulated by their binding partners. Here, we investigate RNA conformational preferences in a non-redundant dataset of 263 protein-RNA complexes to characterize the structural landscape associated with protein recognition. RNA dinucleotide segments are analyzed using seven backbone torsion angles ({delta}1, {varepsilon}1, {zeta}1, 2, {beta}2, {gamma}2, and {delta}2), two glycosidic torsion angles ({chi}1 and {chi}2) and the pseudo-torsion angle . Focusing on dinucleotide steps present in both interface and non-interface regions, we performed density-based clustering using selected backbone torsion angles to identify recurrent conformational states. We identify 28 distinct RNA dinucleotide conformers containing at least ten members each. Among these, eight conformers represent previously unreported nucleotide conformers (NtCs), including the transitional and the non-canonical states AB06, AB07, BB21, BB22, OP32, OP33, IC08 and IC09. Several of these conformers are preferentially enriched at protein-binding interfaces, suggesting their involvement in local conformational adaptation during protein-RNA recognition. The newly identified conformers span transitional A-B geometries, distorted B-like states, open conformations and compact intercalated structures, highlighting the remarkable structural plasticity of RNA in ribonucleoprotein complexes. Overall, this study expands the current understanding of RNA conformational space and provides a refined RNA dinucleotide conformer library for protein-RNA complexes. These findings will facilitate the identification of novel RNA structural motifs and improved RNA structural modeling, docking protein-RNA complexes and deep learning-based prediction frameworks for describing RNA tertiary structures.
Course-based undergraduate research experiences (CUREs) can expand undergraduates' access to research and motivate students to stay in science. Yet, little research has examined how CURE instruction shapes student motivation. We leveraged a motivation-related characterization of non-content talk of 48 CURE and non-CURE instructors to predict the motivation-related outcomes of 462 students. We fit a series of multi-level models (MLM) in which we regressed students' post-course scientific self-efficacy, task values, scientific identity, and science-related intentions onto instructors' self-efficacy and task values-related talk, controlling for students' pre-course levels. We also fit an MLM to explore whether instructors' relationship-building talk (immediacy talk) was associated with students' rapport with their instructor. Instructors' self-efficacy talk did not affect students' self-efficacy, and instructors' immediacy talk had a marginally positive but non-significant association with students' rapport ratings. Instructors' task values talk positively influenced students' scientific identity and some but not all of their task values. Instructors' task values talk also positively influenced students' intentions to pursue a science career, but not graduate education or research careers. Collectively, these results suggest that instructors' task values talk may underpin some of the motivational effects of CURE instruction, but that task values talk need not be limited to CUREs.
Mapping the genetic basis of inter-individual heterogeneity in multifactorial diseases opens the door to mechanistic insights and opportunities for targeted intervention. In Alzheimer's disease (AD), clinical and pathological heterogeneity is well recognized, but genetic dissection is limited by a lack of well-powered cohorts with deep phenotypic characterization. Here, we introduce a polygenic score (PGS) analysis strategy to address these limitations by leveraging the inherent pleiotropy in complex trait genetics. We perform a cross-cohort, cross-trait application of pre-trained PGS, integrating 713 UK Biobank-derived PGS with 36 deep AD phenotypes across 1678 ROSMAP participants. We identify 268 statistically significant (FDR<0.1) associations between 12 prioritized PGS and 36 AD phenotypes. Prioritized PGS include blood lipid measurements, inflammatory biomarkers, and cancer traits; observed AD phenotypes include cognition, amyloid, and tangles. Of the 268 associations, 49 persist with APOE-excluded PGS. Predictive models trained on multiple prioritized PGS outperform the AD PGS or APOE alone for predicting amyloid and cognition. Lastly, our approach identifies six individual-level AD polygenic subtypes supported by distinct pathological patterns. Overall, we combine large-scale biobank resources and deeply-phenotyped cohorts using PGS, reveal genetic features underlying AD heterogeneity, and provide a general model for stratifying heterogeneous disease-focused cohorts using genomics.
Protein labelling by covalent attachment of a specific substrate to a self-labelling protein tag has become a regular in the life sciences. Herein, we report the design of a two-component labelling system, comprised of a non-fluorescent difluorinated xanthene, called F2X, and a HaloTag mutant engineered for targeted reactivity towards F2X. Upon primary covalent locking of the ligand at the canonical aspartate residue, two proximal lysine residues located at the protein surface can undergo nucleophilic aromatic substitution with the F2X core, building a fluorescent rhodamine via triple-covalent fusion. We used a generalizable in silico pipeline for heuristic conformational sampling of covalent protein-ligand complexes to find suitable mutation sites, culminating in the curation of 7 double-lysine HaloTag mutants for targeted in vitro testing. Reaction with the best-performing mutant, HTPL161K_Q165K, is characterized by full protein mass spectrometry, fluorescence polarization fluorescence lifetime, and fluorescence anisotropy and rationalized by computational modelling. We showcase the system in single molecule microscopy, where obviation of post-labelling purification is a prime advantage when targeting recombinant proteins that may not be expressed in larger quantities, and employ F2X in living cells with reduced photobleaching. Lastly, a cell-impermeable version was obtained by means of sulfonation, exclusively targeting extracellularly exposed HTPKK fused to the neuromodulatory G protein-coupled receptor metabotropic glutamate receptor 2.
How do humans store sequences that far exceed working memory capacity? Using visuo-spatial and binary auditory sequences, we previously showed that a Language of Thought (LoT) architecture, in which simple primitives are recursively combined into hierarchical programs, enables efficient storage of structured sequences. Here we ask whether this principle extends to purely ordinal structure: sequences defined by how items repeat and in what order, as in AABBCCAABBCC, independently of their spatial content. Across three experiments, participants reproduced 12-item sequences of spatial locations with various ordinal structures. The minimal description length derived from the LoT model predicted recall accuracy with remarkable precision (r = .96), substantially outperforming Shannon entropy, Lempel-Ziv complexity, chunking models and subjective complexity ratings. Critically, fine-grained analyses of participants' inter-click intervals during reproduction revealed systematic slowdowns at the hierarchical boundaries predicted by the LoT programs, providing a behavioral signature of the underlying mental syntax. These results identify a compact vocabulary of mental primitives, repetition, mirroring, and interleaving, whose composition accounts for the symbolic compression of ordinal structures. For ordinal regularities, human sequence memory operates as a form of program induction, leveraging a domain-general capacity for hierarchical compression to encode complex structured information.
How post-mitotic neurons maintain precise transcription factor (TF) levels throughout life remains a fundamental open question. Here, we challenge the prevailing model of positive autoregulation by demonstrating that UNC-3 (Collier/EBF1-4), a dosage-sensitive TF continuously required for cholinergic motor neuron identity in C. elegans, negatively regulates its own expression. Using genetics, biochemistry, and inducible protein depletion, we show this self-repression occurs directly at the transcriptional level and persists beyond development. CRISPR/Cas9 disruption of negative autoregulation causes motor neuron identity and locomotion defects, establishing its functional necessity. Mechanistically, the UNC-3 DNA-binding domain is required and sufficient for self-repression, with an AlphaFold2 screen implicating chromatin factors as interaction partners. Critically, UNC-3 self-repression is continuously counterbalanced by positive input from the HOX cofactor CEH-20/PBX, revealing a dynamic "balancing act" between opposing regulatory inputs that stabilize TF dosage over time. Mutations in the unc-3 ortholog EBF3 cause a neurodevelopmental syndrome, and disease-associated variants disrupt UNC-3 self-repression, revealing a key molecular mechanism underlying the disorder. We propose that negative autoregulation continuously counteracted by positive input represents a broadly applicable principle for maintaining dosage-sensitive TF expression to secure post-mitotic cell identity.
Despite sharing the same genes and the same environment, individuals often develop substantial phenotypic differences. While this pattern has been documented across diverse species and traits, the processes giving rise to this 'stochastic' or non-shared environmental variation remain unclear. Recent mathematical models of development in which phenotypes are gradually constructed may offer some clues. These models show that imperfect environmental cues can generate striking variation in developmental trajectories and adult phenotypes. At the population level, such imperfect cues produce increasing stability of individual differences across ontogeny (e.g. animal personality) and patterned distributions of mature phenotypes (e.g. normal or skewed) that resemble those observed in real organisms. Our paper synthesizes existing models in which stochastic phenotypic variation arises solely as a by-product of mechanisms missing their phenotypic targets because of imperfect cues. We then link these models to related, but independent, mathematical theory exploring the environmental conditions under which stochastic phenotypic variation is favoured by natural selection. Our integration shows that stochastic sampling is often favoured over classic bet-hedging strategies involving non-plastic generalist or specialist strategies. Our findings provide new directions of research on stochastic sampling as a mechanism for adaptive stochastic variation within and across generations.