Monday, October 21, 2024

Heritability

In a previous post I described “The Fundamental Law of Genetics,” which states that phenotype (observed trait) = genotype (genetic factors) + environment (environmental factors). The quantitative version of the law is $V_P = V_G + V_E$ in which the phenotypic variance in a population equals genotypic variance (phenotypic variation caused by differences in genotype)  plus environmental variance (phenotypic variance caused by differences in environment). Then I cautioned that the equation is a simplification that omits genotype-by-environment interactions.

In a more recent post I described an alternative version of this law as $z$ (phenotypic value) $= G$ (gentoypic value) $+ E$ (residual or environmental value). A phenotypic value is a measurement of a trait like your height, and the genotypic value is the portion attributed to genotype with the remainder being attributed to environment. This value equation can be converted into the variance equation by taking the variance of both sides and then throwing out the covariance term between G and E.

Intuitively, any trait (such as blood pressure) depends on genetic factors (alleles) that you inherit from your parents as well as environmental factors (such as diet and exercise). But how much is genetics and how much is environment? The concept of heritability attempts to estimate the fraction of a trait that is due to genetics (your genotype).

More quantitatively, heritability is the  fraction of total phenotypic variance ($V_P$) that can be explained by genotypic variance ($V_G$):

$ \begin{equation} H^2 = \frac{V_G}{V_P} = \frac{V_G}{V_G + V_E} \tag{1} \end{equation} $

where $H^2$ refers to broad-sense heritability. My students asked whether they needed to take the square root to obtain $H$, but no, $H^2$ not $H$ is defined to be heritability (good question as always).

Estimating heritability is easier said than done, rife with assumptions and caveats. Further, there is a wide gap between the technical definition of heritability, and the common-sense understanding among the general public (e.g. the apple doesn't fall far from the tree) who tend to overestimate heredity and underestimate rearing. Many of these issues are beyond the scope of this post, and here I will focus on one simple approach to estimating heritability by comparing the phenotypes of parent and child.

In the early 1900s, Francis Galton gathered data highlighting the similarities between parents and children. I have reproduced Figure 8.2 from Hartl and Clark which shows a scatter plot of the weight of male offspring on the y-axis with the weight of their father on the x-axis. The fact that this trait is inherited to a certain degree is reflected in the linear relationship between the parental trait value and the offspring trait value, so that the data points can be fit with a line (i.e. linear regression, which Galton helped to invent). More importantly, this line has a positive slope so that higher offspring values correspond to higher parental values.

Approximate reproduction of Figure 8.2 from Hartl and Clark (2007). Mean weight of male offspring of flour beetle (y-axis) plotted against weight of father (x-axis).  
On the one hand, if the heritability was very low for the trait, then the data points would be randomly scattered and the best line to fit the points would be horizontal of slope 0. On the other hand, if the heritability was close to 1, then the children would have the same value of the trait as their parent (in the simplest case), and the points would lie along a line of slope 1. 

The converse is not true; points lying along a line of slope 1 do not imply that the heritability is 1. There could be environmental components which are identical for both parents and children, e.g. both grew up in wealthy households. Environmental covariance refers to related environments (e.g. for parents and offspring) leading to related phenotypes. However, if we assume that the environmental covariance is zero, we can calculate heritability from the linear regression slope.

We can simplify matters by defining the "midparent" which is the average phenotype of the Mother and Father, and plot the offspring phenotypic values (y-axis) against the midparent phenotypic values (x-axis). From statistics, the resulting regression line that best fits these points will have slope $b = \frac{\operatorname{Cov}(x,y)}{\operatorname{Var}(x)}$.

The phenotypic covariance between child ($y$) and midparent ($x$) is a measure of trait resemblance, i.e. how correlated are the two phenotypes (correlation is scaled covariance). As stated above, one can write the phenotypic value $z$ as the sum of the genotypic value $G$ (contribution of genetic factors) and the environmental value $E$ (contribution of environmental factors): $z = G + E$.

In the general case for any two individuals $x$ and $y$, the phenotypic covariance can written as: $\operatorname{Cov}(x,y)  = \operatorname{Cov}(G_x + E_x, G_y + E_y) = \operatorname{Cov} (G_x,G_y) + \operatorname{Cov}(G_x,E_y) + \operatorname{Cov}(E_x,G_y) + \operatorname{Cov}(E_x,E_y)$ where $G_x$ represents the genotypic values and $E_x$ the environmental values for individual $x$. We will assume that the genotype-environment covariances ($\operatorname{Cov}(G,E)$ terms) are 0 so that we can focus on the remaining two terms.

The term $\operatorname{Cov}(E_x,E_y)$ describes the relatedness of the environments of the two. If $x$ and $y$ were random individuals, then we would expect this covariance to be 0, but as mentioned above, for parent and child many environmental conditions are likely to be similar (“I’m raising my kid just like I was raised”) so that the covariance would not be zero for many traits. Assuming this covariance is 0 is the biggest weakness of calculating heritability from parent-offspring regression, but for now we will assume that it is 0 to continue with the blog post using the derivations outlined in Coop (2020), Hartl and Clark (2007), and Lynch and Walsh (1998).

The remaining term is the covariance of genotypic values $\operatorname{Cov} (G_x,G_y)$. As discussed in the previous post on dominance and epistasis, each genotypic value can be broken up into additive, dominance and epistatic terms: $G_x = G_{xA} + G_{xD} + G_{xS}$. We can then write $\operatorname{Cov} (G_x,G_y) = \operatorname{Cov}(G_{xA},G_{yA}) + Cov(G_{xD},G_{yD}) + \operatorname{Cov}(G_{xS},G_{yS}) +$ cross terms that are 0. For parent and offspring, $\operatorname{Cov}(G_{xD},G_{yD}) = 0$, and most of the epistasis terms are also 0 or presumed to be small (see Table 7.2 in Chapter 7, Lynch and Walsh for details). This leaves us with $\operatorname{Cov} (G_x,G_y) \sim \operatorname{Cov}(G_{xA},G_{yA})$. As a reminder, the additive term is from the linear model of summing the contributions of individual alleles that influence a trait, whereas dominance and epistasis represent nonadditive genetic contributions.

In close relatives, one can calculate the covariance of the additive genotypic values by following the descent of alleles to the respective individuals. Given a locus, an offspring inherits one allele from one parent and one allele from the other parent. The additive covariance between parent and child is 0 for the allele not inherited from the parent (correlation = 0), and is $\frac{V_A}{2}$ for the inherited allele, in which $V_A$ is the additive genotypic variance for the locus with the $\frac{1}{2}$ arising from that a locus consists of two alleles. Thus, for the locus we have $\operatorname{Cov}(G_{xA},G_{yA}) = 0 + \frac{V_A}{2} = \frac{V_A}{2}$. Similar reasoning applies to calculating the additive genotypic covariance between offspring and midparent.

The above is an example of the concept of identity-by-descent (IBD) which refers to shared, completely identical DNA segments between two or more individuals inherited from a common ancestor. A parent shares 50% of DNA with a child through IBD, and so 50% of the offspring's alleles are identical to the parent's through direct descent.

As described above, we can write the slope $b$ of the midparent-offspring regression as the covariance divided by the variance of the midparent which is $Var(x) = \frac{V_P}{2}$ because the midparent is the phenotypic average of the two parents. This leads to the following equation:

$ \begin{equation} b = \frac{\operatorname{Cov}(x,y)}{\operatorname{Var}(x)} \sim \frac{\operatorname{Cov}(G_{xA},G_{yA})}{\operatorname{Var}(x)} = \frac{V_A}{V_P} = h^2 \tag{2} \end{equation} $

The expression $h^2 = \frac{V_A}{V_P}$ is referred to as narrow-sense heritability which consists only of the additive component of the genotypic variance (omitting the dominance and epistatic terms) divided by the total phenotypic variance, and hence it is always less than or equal to broad-sense heritability. However, both theory and empirical data indicate that $h^2$ is close to $H^2$ (Benjamin, et al. 2024). 

Thus, we can calculate the narrow-sense heritability by performing linear regression of offspring phenotypic values against parental phenotypic values for a trait. The main limitation is the assumption that the environmental covariance of parent and offspring is 0 ($\operatorname{Cov}(E_x,E_y) = 0$). If the covariance is non-zero, then the estimate of $h^2$ will be upwardly biased because there is an additional positive term in the numerator of Eq 2 that shouldn't be there. Because of this suspect assumption, researchers in the past have turned to a different method using twins to estimate heritability, the subject of a future post.

TL;DR: There is broad-sense and narrow-sense heritability which describe the fraction of a trait that is attributed to genetic factors. The latter can be estimated by parent-offspring regression, but with the key assumption that environmental covariance is 0. However, it is very unlikely that the parental environment is uncorrelated with the offspring environment. The apple and the tree grow on the same patch of land. 

No comments:

Post a Comment

ChatGPT identified 142 statistical "issues" in problematic paper

Today the Bayesian statistician Andrew Gelman posted a rather pointed critique of a paper on his blog stating: "Wow! This paper is an ...