Research

The research in our lab revolves around how evolutionary forces such as mutation, natural selection, and population history shape genetic variation, primarily at the population level.

We are particularly interested in the impacts on human traits and diseases and develop population-genetic theory, statistical methods, and computational tools to address these questions.

Mutation rates · rare genetic variation

Mutation rates and rare genetic variation

New mutations are the source of all genetic variation within populations and ultimately all evolutionary change. Different sites in the genome differ by orders of magnitude in their germline mutation rates, and these differences dominate the distribution of rare variation observed in sequencing studies.

We develop theory and methods to understand mutation rates and to better utilize rare variation in genetic analyses. We previously worked on Roulette, which provides basepair-resolution predictions across the human genome, and current efforts are focused on improving accuracy and providing a complete description of hypermutable elements in the human genome.

We also developed population-genetic models for how recurrent mutations in large samples and at high mutation-rate sites affect allele frequencies, and applied these models to improve the identification of mutational processes, estimate mutation-rate distributions in hypermutable regions, and fit recent demographic history.

While an overabundance of variation relative to model predictions can indicate enhanced mutability, depletions of genetic variation indicate selective constraint and possible disease contributions. Highly mutable sites that remain invariant in enormous sequencing cohorts provide strong evidence for selection.

Residual mutation-rate variance across validation datasets for Roulette and prior models
Rare-variant frequency inference identifies an extreme mutation-rate distribution within snRNA genes

Natural selection · complex traits

Natural Selection and the Genetic Architecture of Complex Traits

Complex-trait variation arises from the interplay of mutation, genetic drift, and natural selection. Biobank GWAS enable analyses of genetic architecture across large collections of traits and diseases. To explain the patterns uncovered by these studies, we modeled the evolutionary origins of genetic architecture and used these insights to infer selection from GWAS data.

The study of natural selection using GWAS largely relies on single-locus models, but many properties of phenotypic selection, mutational architecture, and genetic interactions are not identifiable from single-locus models. We work on the development of two-locus models as a flexible and tractable way forward. Variant pairs provide a starting point for learning how variant effects compose on haplotypes.

These models connect patterns of signed effects among linked variants to features of trait architecture including the genomic distribution of causal mutations, pleiotropy, polygenicity, linkage disequilibrium, and recombination.

Neutral, directional, single-trait stabilizing, and pleiotropic stabilizing selection models
GWAS effect sizes and risk-allele frequencies for type 2 diabetes, colored by estimated stabilizing selection
Evolution of haplotype frequencies under neutral and negative selection

Functional effects · selection

Mapping Functional Effects to Selection

Selective constraint, the propensity of the genome to remain unchanged across evolutionary time and within populations, is a precise indicator of broad functional importance. Expanded phylogenetic and human sequencing data, together with functional genomic measurements at base-pair resolution and in individual cell types, provide an opportunity to characterize the biology mediating constraint.

High-throughput experimental approaches and computational predictors now allow variant effects to be characterized at scale. Under negative selection, human polymorphism provides a readout of fitness effects. Our lab is developing population-genetic inference techniques that efficiently and flexibly map predictor scores to selection coefficients, allowing measurements and predictions from different molecular contexts to be compared on a common fitness scale.

A major goal of this effort is to identify when and how functional assays, sequence-based predictors, and population-genetic evidence should be combined into a joint framework.

Workflow connecting experimental and computational effect predictions to population data and selection