next on phyloseminar.org
To attend a seminar, please visit the livestream portion of our YouTube channel.
Advances in Mutation and Selection Models
Modeling site-and-branch heterogeneity in phylogenetic datasets
Due to widely varying mutation and selection effects, realistic phylogenetic models for protein sequences from diverse taxa require that amino acid frequencies vary over both the sites and branches. Failure to model this compositional heterogeneity can lead to phylogenetic artefacts. However, the computational cost of phylogenetic inference with models accounting for compositional heterogeneity can be prohibitive. We present the modelling frameworks that we are using to adjust for this variation in frequencies of amino acids over sites in an alignment and over the tree. These include composite likelihood approaches for frequency estimation and the GFmix modeling framework which allows groups of amino acids to vary over branches in a tree. We argue that the use of mixture models ameliorates problems of over-parameterization and demonstrate through simulation and empirical examples the benefits of the framework.
Attention on Evolution: Foundation Models for Molecular Adaptation and Epistasis
Detecting molecular adaptation and allosteric co-evolution across macroscopic phylogenies is fundamental to evolutionary genomics and biology. However, workhorse method -- classical codon models -- require numerical optimization which scales poorly and become sensitive to data quality for modern datasets, rendering whole-genome sweeps computationally very expensive and noisy. I'll discuss HyphAeon, a phylogenetic nano-scale foundation model that replaces iterative likelihood fitting with an axial self-attention architecture. HyphAeon directly integrates evolutionary tree topologies into row attention through continuous-time Markov transition kernels and tree-scaled rotary embeddings, while employing block-diagonal channel disentanglement to prevent synonymous rates from confounding selective pressure.
Operating up to 100,000 × faster than classical maximum likelihood (like HyPhy), HyphAeon screens whole mammalian (VGP scale) and avian genomes in minutes on standard consumer hardware. Beyond accelerating selection inference, training on phylogenetic alignments unlocks rich, unprompted emergent behaviors. The model’s internal representations capture multi-site epistatic co-evolution, delineating collective sectors that map functional allosteric networks. Furthermore, by evaluating lineage-specific episodic selection across thousands of homologous genes simultaneously, HyphAeon powers genome-wide phenotype-to-genotype association mapping—uncovering convergent molecular adaptation underlying complex ecological, dietary, and life-history traits. Simultaneously, digital deep mutational scanning quantifies sitewise genetic plasticity in seconds, transforming comparative genomics from a descriptive statistical pipeline into an ultra-fast, predictive engine for molecular evolution.
Unlocking the rules of proteome evolution
Why does a cell produce a particular amount of protein? Why does a protein-coding sequence use a particular set of nucleotides? These questions are central to understanding the evolution of the proteome across species. Here, I will present work on two key aspects of proteome evolution using models that decompose the influences of mutation and natural selection to address these questions. First, the non-uniform usage of synonymous codons, or codon usage bias (CUB), is a universal feature of genomes, but varies in direction and magnitude across species. Using a population genetics model assuming selection-mutation-drift equilibrium, we quantified CUB across 327 budding yeast species. We linked variation in natural selection on synonymous codon usage with changes to the tRNA pool, consistent with selection for efficient translation. Furthermore, we find evidence that changes to the tRNA pool reflect changes in genome-wide nucleotide content. This suggests a model in which mutation bias shapes the selectively favored synonym via the tRNA pool, a key factor determining codon-specific translation efficiency. Second, changes to gene expression are a major driver of phenotypic variation across species. Gene expression results from multiple processes, from transcription through post-translational regulation, requiring evolutionary models that capture their mechanistic relationships. The regulation of mRNA and protein abundances is well studied, but less is known about the evolutionary processes that shape their relationship. We derived a new phylogenetic model and applied it to mRNA and protein abundance data across ten mammalian species. Our analyses reveal strong stabilizing selection on protein abundances over macroevolutionary time, that mutations affecting mRNA abundance minimally affect protein abundance, that mRNA abundances evolve under selection to track protein abundances, and that mRNA abundances adapt faster than protein abundances due to greater mutational opportunity. Together, these projects show how decomposing mutation and selection can reveal the evolutionary rules governing the proteome, from codon choice to the coevolution of gene expression.