Today the Bayesian statistician Andrew Gelman posted a rather pointed critique of a paper on his blog stating: "Wow! This paper is an absolute clinic in bad quantitative social science research. I’ll leave it as an exercise for the reader to count up all the problems."
Well, ChatGPT-5.6 Sol (Pro) compiled a list of 142 "issues" with the paper far exceeding my paltry efforts. For comparison, Claude (Fable 5 Max) found 16, and Gemini 3.1 Pro was satisfied with 7. Admittedly the ChatGPT list contains a fair amount of repetition, i.e. same error but described in different terms, and nit-picking.
After prompting from me, ChatGPT grouped the 142 problems into 9 categories.
After more prompting, it mapped common statistical fallacy/bias names on to its 9 categories.
Finally, ChatGPT formulated its own study design. which it contrasted with the methodology in the paper concluding:
"Thus, the paper compared selected deceased athletes with an aggregate population benchmark, whereas a rigorous study would compare defined groups of living and deceased athletes over observed person-time properly handle censoring and truncation, balance confounders, and estimate standardized differences in mortality or survival. The latter could support carefully qualified sport-specific associations; the paper’s method cannot validly estimate how many years a sport adds to or subtracts from life."
I am not qualified to assess the responses from ChatGPT, but I trust that they would be a good starting point for an expert human referee reviewing the paper whose first task would be whittling down the list to remove minor and duplicated errors.
At the very least, it is a valuable statistics lesson for me from my AI statistics tutor.
A complex trait is defined to be a non-Mendelian trait, which means that the trait does not obey the rules and assumptions of Mendelian genetics such as 1) a single gene determines the trait, 2) there are only dominant and recessive alleles for each gene, and 3) the phenotype is discrete and not continuous.
Complex traits are multifactorial and quantitative, violating Mendel’s assumptions. In a multifactorial trait, multiple genetic and environmental factors contribute to the phenotype. A quantitative trait possesses not two or three possible phenotypes (e.g. round or wrinkled) but an infinite number along a continuous scale (e.g. slightly wrinkled).
In an earlier post, I stated that the fundamental law of genetics is that Phenotype = Genotype + Environment, and that heritability is the portion of a trait (phenotype) that can be attributed to genotype (genetics) and not environment. A classic Mendelian trait possesses heritability = 1 (gene determines the trait). In the real world, human complex traits are multifactorial and hence have heritability much less than 1 and often close to 0 (i.e. environmental factors contribute more to trait than genetic factors).
How can we identify the genetic factors (i.e. genes) that influence a given trait of interest? For example, blood pressure is a complex trait possessing multiple genetic and environmental factors. The estimated heritability is approximately 0.25 (SNP) to 0.5 (twin studies). What genetic variants (i.e. gene alleles) contribute to high blood pressure? Identifying the relevant genes can help predict risk of hypertension from one’s genome sequence (even before the phenotype manifests itself), and can serve as targets for therapeutic treatments.
One approach is linkage analysis which follows the segregation of genetic markers (i.e. SNPs) and the phenotype in family trees. Co-segregation indicates that the genetic factor contributing to the trait and the marker are close together (linked) on the same chromosome (i.e. not likely to recombine and become separated). This method works best for Mendelian traits in which a single gene “determines” the trait with little to no influence from other genetic and environmental factors so that the genetic signal is strong.
As a reminder, a SNP (single nucleotide polymorphism) is defined as a variation at a single position in a DNA sequence among individuals of a species. SNPs play a central role in human genetics because they can serve as genetic markers. These markers are “mile posts” of chromosomal (genome) locations (loci) analogous to a map. Certain SNPs are close to or within certain genes.
The second, more commonly-used approach of identifying phenotype-influencing genetic factors is a Genome-Wide Association Study (GWAS), which is a powerful observational study design to pinpoint genetic variants that are statistically associated with a particular trait or disease. GWAS examines a large set of genetic markers (i.e. SNPs) spanning the entire genome so that it isn't restricted to a single gene or a small set of candidate genes. Importantly, association does not prove causation; it simply means that a particular variant is found more often in individuals with the trait than in those without. In this manner, one can identify regions of the human genome that associate with the trait and focus in on genes residing in that region.
For example, the "T" allele of a certain SNP (rs16998073) is associated with the trait of high blood pressure, and so people containing the T nucleotide at a particular location in the human genome (in the vicinity of 4q21: chromosome 4, long arm (q), band 21) are more likely to have high blood pressure than people with a different nucleotide at that position in the genome (link). This SNP marker happens to be close to some gene(s) presumably involved in blood pressure control. Further work is needed to demonstrate a causal link with a particular gene and a particular variant (allele) of the gene.
Association between genetic marker and trait arises because of linkage disequilibrium (LD), which refers to the non-random association of alleles at two or more different loci (locations on a chromosome) in a given population. In simple terms, an allele (variant) at one location is disproportionately likely to be found together with a certain allele at another location on the same chromosome. Not surprisingly, the main cause of LD is linkage (physical proximity). Alleles at loci that are physically close together on the same chromosome tend to be inherited together because recombination events (crossing over during meiosis that separate the alleles on to different chromosomes) are less likely to occur between them. Other causes of linkage disequilibrium include epistasis, population admixture, genetic drift and recent mutation.
A hypothetical causal mutation that influences a trait can be in linkage disequilibrium with a nearby SNP genetic marker so that the marker associates with the trait. We need to genotype (determine alleles) of many SNP markers in a large population to find out which ones associate with the trait, and thus localize the hypothetical causal mutations to particular genes near the SNP "hits". Remember there are only about 20,000 genes in the human genome.
The basic steps for performing a GWAS are as follows:
1. Recruit study population. Decide on a (disease) trait to investigate and gather hundreds of thousands, even millions of unrelated individuals. They are divided into cases (individuals with the trait) and controls (without the trait). Why so many? Most common diseases are influenced by many genes (polygenic), each with a very small effect. A huge sample size is required to have enough statistical power to detect these subtle effects.
2. Genotyping. DNA samples (usually from blood or saliva) are collected from participants, and analyzed by SNP arrays or more recently by whole genome sequencing (WGS). This provides information on the particular allele (nucleotide) at each SNP being assessed in the GWAS for every person. Typically hundreds of thousands to millions of SNPs scattered across the human genome are genotyped for every subject.
3. Statistical analysis. Each SNP is tested for association with the trait over the whole study population. For illustrative purposes let’s assume that the trait is discrete and binary, e.g. disease or no disease. Then for each SNP (assuming 2 alleles, which are referred to as genotype 1 and genotype 2), one can construct the following 2x2 contingency table with $a, b, c, d$ representing the numbers of people in each category i.e. genotype 1 with disease.
Genotype 1 would be one allele at a particular SNP (e.g. "A" at SNP rs16998073), and Genotype 2 would be another allele at that SNP (e.g. "T"). The variables $a$, $b$, $c$, and $d$ are the number of people with a particular genotype and phenotype that sums up to total study population.
From this table one can calculate the odds ratio (OR) which is a measure of effect size (i.e. how much the particular SNP will affect the trait). It is the ratio of the odds of an event occurring in one group relative to the odds of it occurring in another group (odds is the probability of disease divided by probability of no disease). Thus, the odds-ratio for the contingency table above is $OR = \frac{a}{b} / \frac{c}{d} = \frac{ad}{bc}$. If $OR = 1$, no association between genotype and outcome; if
$OR > 1$, SNP genotype (allele) is associated with higher odds of the outcome; and if $OR < 1$, genotype is associated with lower odds of the outcome.
One can then perform the chi-square significance test to assess the statistical significance of the OR deviating from 1, which is summarized as a p-value for the particular genotype (SNP allele). Thus, one obtains a p-value for every SNP that is tested in the GWAS for association with the trait of interest.
An alternative (and more common) approach to GWAS statistical analysis is logistic (for binary traits) or linear (for continuous traits) regression in which a linear model describes the contribution of each SNP to the trait. A big advantage of using regression is that adjustments can be made for potential confounders (e.g., age, sex, ancestry differences).
4. Interpreting and Displaying Results. The final results of the statistical analysis (SNP p-values) are visualized in a Manhattan plot, which is a scatter plot displaying SNP associations across chromosomes. Each data point represents one SNP with its chromosome position along the x-axis (i.e. in which chromosome is it located), and the y-axis is the p-value for the association with the trait from the statistical analysis. SNPs that surpass a stringent significance threshold (typically a p-value of <5 × 10⁻⁸ which accounts for multiple testing of hundreds of thousands of SNPs at the same time) indicate genomic regions associated with the trait. Usually a group of SNPs in the same region, presumably near a causal mutation/variant, will rise above the background. It is called a Manhattan plot because of the visual appearance of the Manhattan skyline with high-scoring hits appearing as skyscrapers.
Each dot represent one of numerous (hundreds of thousands of SNPs). They are color-coded according to their chromosomal location on the x-axis. The y-axis plots $-\log_{10}(\text{p-value})$ of the SNP association with the trait as measured by chi-square or linear/logistic regression. The red line indicates the threshold above which a SNP p-value is considered statistically significant ("hit"). In this example, a cluster of SNPs near the gene TCF7L2 clear the threshold implicating this gene for the trait (type 2 diabetes).
It is important to acknowledge the limitations of GWAS. First, it usually identifies common genetic variants with modest effects. Rare variants with larger effects may not be detected because they are so rare, and variants with very small effects may also be missed. Second, population stratification (genetic ancestry differences) can lead to false positives if the confounding is not properly accounted for. Certain variants may be prevalent in certain subpopulations, and so may associate with traits that are enriched in the subpopulations for environmental (e.g. cultural) reasons (e.g. chopstick gene). Third, GWAS requires very large sample sizes to achieve adequate statistical power. Finally as mentioned above, association does not imply causality. There may be confounding factors so that the SNP association with the trait is spurious and not the result of tightly-linked (nearby) causal genetic factors.
Despite these limitations, there has been a flood of new information over the last 20 years. It is no exaggeration to say that GWAS has ushered in a revolution in human genetics identifying numerous genetic factors associated with countless traits ranging from coronary artery disease to Alzheimer’s to type 2 diabetes. One example among many is how GWAS pinpointed CFH and ARMS2/HTRA1 as major risk genes for age-related macular degeneration (AMD). These findings have led to new biological understanding as GWAS hits revealed unexpected genes and pathways such as PCSK9, which later became the target of cholesterol-lowering drugs. Indeed, the genetic associations can aid in the discovery and validation of drug targets. Last but not least, GWAS is a necessary step for constructing Polygenic Risk Scores (PRS, also known as polygenic score or PGS) which uses SNPs to estimate the genetic risk for a disease trait. This is sort of the reverse of GWAS which starts with the trait and then identifies genetic factors that contribute to the trait. These associated SNPs can be used in a PRS to predict the likelihood of the trait in new individuals from a different population from their genotypes. PRS have already entered the clinic, and they have provided significant additional information beyond standard clinical risk scores which are derived from lab tests.
TL;DR: GWAS is a powerful tool for identifying genomic regions that associate with a particular trait. This has allowed researchers to pinpoint specific genes that influence the trait.
I asked Gemini (Pro 2.5) "Under what conditions is lidar better than cameras for self-driving cars?"
Of course. Lidar is better than cameras for self-driving cars under specific conditions where its core strengths—direct distance measurement and immunity to lighting issues—outweigh a camera’s advantages in interpretation and cost.
The fundamental difference to remember is:
Lidar is an active sensor: It emits its own light (lasers) and measures the reflection. This makes it a direct 3D measuring device.
Cameras are passive sensors: They rely on ambient light (from the sun or streetlights) to capture 2D images. They infer depth and distance using complex software.
Here are the key conditions where lidar is superior:
1. Low Light and Complete Darkness
This is lidar’s most significant advantage.
Why Lidar is Better: Since lidar provides its own light source, its performance is nearly identical whether it’s noon or a moonless midnight. It can build a precise 3D map of its surroundings, detecting pedestrians, curbs, and other vehicles with high accuracy, regardless of the time of day.
Camera Limitation: Standard cameras struggle in the dark. While they can be improved with software (like long exposure or ISO gain), this often introduces image noise, motion blur, and reduced detail, making it harder to reliably detect unlit objects.
Example: A person wearing dark clothing walking on an unlit road at night is a classic nightmare scenario for a camera-only system. Lidar will detect them as a clear object at a precise distance.
2. Direct Sun and Sudden Glare
Cameras can be easily blinded, just like the human eye.
Why Lidar is Better: Lidar operates on a very specific, narrow wavelength of light (usually in the infrared spectrum) and uses filters to block out all other light, including sunlight. It is therefore almost completely immune to glare from the sun rising or setting, or reflections off wet roads or other cars.
Camera Limitation: Driving directly into a low sun can “white out” a camera’s sensor, making it temporarily blind to everything in that direction, including brake lights or a stopped car.
Example: Exiting a dark tunnel into bright daylight. A camera’s aperture needs time to adjust, during which its perception is compromised. Lidar’s performance is unaffected.
3. When Unambiguous, Precise Distance is Critical
Lidar’s primary function is to measure distance directly and accurately.
Why Lidar is Better: Lidar measures the “time-of-flight” of a laser pulse to calculate distance. This is a direct physical measurement, not an inference. This allows the car to know with centimeter-level accuracy exactly how far away an object is and its precise shape. This is crucial for calculating “time-to-collision” and making split-second decisions.
Camera Limitation: Cameras must estimate distance. Monocular (single) cameras do this by analyzing the apparent size of known objects (e.g., “that car looks small, so it must be far away”), which can be fooled. Stereo (dual) cameras are better, but they still struggle with textureless surfaces (like a white truck trailer) and have lower accuracy at longer ranges.
Example: A flatbed truck is carrying long pipes that extend past the end of the bed. A camera might only identify the truck bed, failing to correctly perceive the distance to the pipe tips. Lidar’s point cloud would clearly map the pipes as the nearest obstacle, preventing the car from driving underneath them.
4. Detecting Unfamiliar or Unusually Shaped Objects
Lidar creates a “geometric truth” model of the world.
Why Lidar is Better: Lidar doesn’t need to know what an object is to know that it’s there. It simply maps a collection of points in 3D space. A fallen tree, a construction barrier, or a deer on the road are all detected as solid objects to be avoided, even if the AI has never been trained on that specific thing.
Camera Limitation: A camera-based system relies heavily on its training data to classify objects. An object it has never seen before might be misidentified or missed entirely.
Example: Debris on the highway after an accident. Lidar will see a cluster of points to navigate around, whereas a camera’s AI might struggle to classify the mangled objects.
5. In Some Adverse Weather (with a caveat)
Lidar can often see through certain types of weather better than a camera.
Why Lidar is Better: The infrared laser used by lidar can penetrate through light fog, dust, and rain more effectively than visible light can. It provides a more usable signal in conditions where a camera image would be a blurry or obscured mess.
Camera Limitation: A camera lens covered in raindrops or looking into thick fog sees very little, severely degrading its perception capabilities.
The Caveat: Lidar is not immune to weather. Heavy snow, dense fog, or torrential rain can scatter the laser beams, creating noise and reducing its effective range. However, it often degrades more gracefully than a camera, still providing some useful data when a camera provides none.
Summary Table: Lidar vs. Camera
Condition
Lidar (Better)
Camera (Worse)
Darkness/Night
Unaffected; provides its own light.
Severely limited; relies on ambient light.
Sun Glare
Immune; uses specific laser wavelength.
Can be blinded, causing temporary failure.
Distance Accuracy
Direct, centimeter-level measurement.
Inferred/estimated, less reliable.
Object Shape
Precisely maps 3D geometry.
Can misjudge complex shapes and overhangs.
Novel Objects
Detects anything with physical form.
May fail to identify untrained objects.
Light Fog/Dust
Can often penetrate for a clearer view.
Vision is easily obscured.
Conclusion: The Power of Sensor Fusion
It’s important to note that the debate is not strictly about one being “better” than the other overall. The safest and most robust self-driving systems (like those from Waymo and Cruise) use sensor fusion.
They combine Lidar (for geometric truth), Cameras (for context, like reading signs and traffic lights), and Radar (for excellent velocity detection, even in heavy rain/snow) to build a comprehensive, redundant model of the world. Each sensor covers for the weaknesses of the others, creating a system that is far safer than any single sensor could be alone.
I asked Gemini to write a summary of the status of solid-state batteries in electrical vehicles (EVs):
The Basics
Core Technology: Replaces the flammable liquid electrolyte in current lithium-ion batteries with a solid material (ceramic, polymer, etc.).
The Goal: To create a safer, more powerful, and faster-charging battery for EVs.
The Promised Benefits
Vastly Improved Safety: Eliminates the risk of battery fires by removing flammable liquid.
Higher Energy Density: Could enable 50-100% more range or significantly smaller and lighter battery packs.
Faster Charging: Aims for a 10-15 minute charge from 10% to 80%.
Longer Lifespan: More resistant to degradation, potentially lasting for over 500,000 miles.
The Major Hurdles
Manufacturing at Scale: This is the #1 challenge. Moving from tiny lab-made cells to mass-producing large, reliable, and affordable EV battery packs is extremely difficult.
Material Science Issues:
Maintaining good contact between the solid layers as they expand and contract.
Ensuring ions can move quickly through the solid material, especially in cold weather.
Still a risk of "dendrite" growth (lithium spikes) that can short-circuit the cell.
High Cost: Current materials and manufacturing methods are far more expensive than mature lithium-ion technology.
Key Players & Current Progress
Toyota: A research leader, targeting a limited commercial launch around 2027-2028.
QuantumScape (backed by VW) & Solid Power (backed by Ford, BMW): Leading startups that are now delivering early "A-sample" prototypes to automakers for testing and validation.
NIO: Already delivering a car with a "semi-solid-state" battery, which is a transitional technology that offers some benefits but is not a true all-solid-state battery.
Battery Giants (CATL, Samsung, LG): All have active R&D programs to ensure they are not left behind.
Realistic Timeline & Conclusion
Current Status: Advanced R&D and pilot-production phase. Not yet available in mass-market cars.
2027-2030: The earliest we might see the first true solid-state batteries in a few niche, high-end EVs, but volumes will be very low.
2030s and Beyond: The decade where solid-state batteries have the potential to become mainstream, but only if they can overcome the immense manufacturing and cost challenges to compete with conventional lithium-ion batteries.
The summary of the summary (TL;DR) is that solid-state batteries possess several important advantages over liquid-electrolyte (current generation) batteries in terms of higher energy density, faster charging, safety, and potentially longer lifespan. The main disadvantage is cost, but many battery and car companies are working on the technology to make it cheaper and mass production will also help. Gemini predicts that cars with solid-state batteries "have the potential" to become mainstream in "2030s and beyond".
I made this request because the following article in Electrek (shared on Bluesky by Dean Baker) described how BYD (surprise, surprise) was "planning to launch its first vehicles powered by the new batteries in 2027... Between 2027 and 2029, production will be limited during the first two years. However, in 2030, BYD plans to begin mass production." This fits the timetable outlined above.
What about Tesla? According to an article in Citywire “Tesla has decided not to go down the solid-state battery route and has been focusing on improving its lithium-ion (liquid electrolyte) technology”
That being said the article goes on to state that "Historically, Tesla has relied on outside providers such as Panasonic, CATL and LG Energy Solution to provide it with batteries." So Tesla can purchase solid-state batteries from these companies if necessary. Overall Tesla has adopted a hybrid approach employing both in-house battery production, as well as out-sourcing batteries from major global suppliers.
In the not-too-distant future, everyone will have their DNA genome sequenced. One day you will receive an email from the sequencing company with a link to a ~6.4 GB text file containing roughly 6.4 billion characters. Instead of 26 different letters, the text will consist of only 4 characters: A, G, C, and T representing the 4 nucleotides composing DNA. The sequence will be divided into 46 long contiguous text strings each representing a chromosome (single DNA molecule). Pedagogically, each chromosome sequence can be thought of as a chapter in a book, the very special Book of Me.
The chapters will come in pairs representing the homologous chromosomes, e.g. maternal and paternal chromosome 1, which would represent the closely related Chapter 1M and Chapter 1P. There are approximately 20,000 paragraphs spread across each set of 23 (maternal or paternal) chapters (or ~40,000 paragraphs in total). Like the chapters, the paragraphs come in pairs with a particular paragraph in Chapter 1M, being nearly identical to the corresponding paragraph in Chapter 1P.
You can compare your Book of Me with that of your friend or sibling. In a previous post, I mentioned that you share 50% of your DNA (identity-by-descent, IBD) with a full sibling. Does that mean that 50% of your sequence will be identical to the corresponding sequence in your brother or sister? No.
Instead, you find that your sequence is approximately 99.95% identical to that of your sibling (excluding differences between X and Y sex chromosomes). Perhaps, even more surprisingly, the sequence in your Book of Me will be 99.9% identical with that of a random stranger, no matter what the race or ethnicity. Between any two random individuals there is roughly one character difference every 1000 characters (base pairs; because DNA is double-stranded we speak of a single position on DNA as comprising a base pair, which are the two nucleotides on the complementary strands hydrogen-bonded to each other).
A SNP is a single nucleotide (character) difference between the sequences of two individuals. Note we show the sequence from one of the two DNA strands in the 5' to 3' direction.
This sequence difference is referred to as a single nucleotide variant (SNV) or alternatively a single nucleotide polymorphism (SNP), which is defined as a variation at a single position in a DNA sequence among individuals of a species. Technically speaking, a SNP occurs when a single nucleotide (A, T, C, or G) is altered at a specific location in the genome that is found in at least 1% of the population, i.e. it is relatively common. A SNV refers to any single nucleotide change, regardless of its frequency, i.e. it can be rare. I will use the terms SNP and SNV interchangeably.
Aside from SNPs, there will be other types of sequence differences between two individuals representing other sources of genetic variation. At the largest scale, there are chromosome level changes such as chromosome deletions, inversions, duplications, etc. that can span millions of base pairs or even a whole chromosome (aneuplody = having extra or missing chromosomes such as trisomy 21 which is Down Syndrome). At the smallest scale is the aforementioned SNP at the single base pair level.
Table 12.1 from Genetics: From Genes To Genomes (7th Edition)
In between, are the InDels, which are short (1-100bp) insertions or deletions, and the simple sequence repeats which are short (1-10bp) sequences repeated up to 100 times. There are also larger (thousands of base pairs) DNA segments that can encompass whole genes that are repeated called Copy Number Variants. Of these different classes of genetic variants, SNPs are by far the most prevalent at roughly 1 every 1000 base pairs. The others are 10 to 100 fold less frequent. Thus, when we speak of genetic variation we are mainly talking about SNPs.
SNPs are typically biallelic (two alleles), a major allele (more common) and a minor allele (less common). In total there are 4 possible alleles (the 4 possible nucleotides), but not enough time has passed since the first modern humans (roughly 100,000 years) for random drift to sample all possible nucleotides at each position. The polymorphism is usually reported as the minor allele frequency (MAF). The smaller the MAF, the less polymorphic, i.e. most people in the population will possess the major allele and thus be identical at the locus.
One can consider the minor allele to be a mutant allele in a certain sense, and ask whether this mutant allele has a corresponding mutant phenotype. The larger mutations, i.e. chromosomal rearrangements, are expected to have a significant phenotypic impact because they can alter the expression of one or more genes. A SNP, on the other hand, perturbs only a single base pair, and typically the minor allele does not produce a phenotype different from the wild-type (major allele) base pair. The principal reason is that the SNP locus (position) is typically in the intergenic region (between genes) or in introns, neither of which may have a functional role. Indeed the protein coding region of genes represents only 1.5 - 2% of the human genome. Even if located within the gene coding region, the SNP may not alter the protein sequence (e.g. silent mutation), or may make a conservative amino acid substitution that has minimal effect on protein regulation or function.
SNPs play a central role in human genetics because they can serve as genetic markers. Although any given SNP (technically speaking, SNP allele) may not be the causal mutation of some human trait of interest, SNPs can be in close proximity (linked) to the causal mutation, which could be a different SNP or combination of SNPs or one of the other types of genetic variants (e.g. InDel). Because of this linkage, the SNP will associate with the mutant phenotype in a population, and segregate with the phenotype in a pedigree (recombination is unlikely to separate a particular SNP allele from the causal mutation). Genotyping (i.e. sequencing) the SNP to determine the allele (major or minor) in an individual can thus indicate the presence of the linked causal mutation and phenotype.
As the main source of genetic variation, SNPs are also critical for characterizing human ancestry and diversity. The first sequenced human genome ultimately became known as the reference genome. Subsequent genome sequences could be compared to this reference sequence as well as to each other. As mentioned above, any two genomes differ at approximately 1 in 1000 base pairs, but more related individuals are expected to have fewer differences than less related. Three of the earliest sequenced genomes were those of James Watson, Craig Venter, and an Asian man with initials YK (also involved in the Human Genome Project). Each of the genomes were compared to the reference genome and to each other with the figure below showing the number of SNP differences compared to the reference genome.
Single Nucleotide Variations in each of three genomes relative to reference genome. Each circle contains ~3 million single nucleotide differences from the reference genome (Genetics: From Genes To Genomes, 7th Edition). Of these, approximately 1 million are found in one individual but not the other two, and another 1 million are shared by all 3 (but not in reference). Then there another 500 thousand found in 2 but not in the third (or reference genome).
One surprise was that the number of SNPs shared between Venter and Watson, two Caucasians, was 1.7 million, whereas Venter and YH shared 1.6M, and Watson and YH shared 1.65 million. A naive expectation was that Watson and Venter would share far more SNPs than they shared with the Asian YH. Instead the differences appeared to be small (50,000-100,000) especially compared to the large number of SNPs shared between YH and the other two, roughly 1.6 million each.
This observation hints that the so-called races may not be so different at the genetic level. We can formalize the insight using Wright’s fixation index ($F_{ST}$), which is a measure of genetic differentiation between subpopulations relative to the total population. According to Gusev, the classic definition of $F_{ST}$ for a single site is "the correlation between gametes chosen randomly from within the same subpopulation relative to the [total] population." It quantifies the proportion of genetic variance that is due to differences between subpopulations compared to within subpopulations.
Thanks to the automated DNA sequencing revolution, hundreds of thousands of human genomes have now been sequenced. As a result, researchers have been able to calculate $F_{ST}$ for various pairs of subpopulations. Among geographically defined groups (e.g. East Asians, sub-Saharan Africans, Europeans), $F_{ST}$ between any two was estimated to be in the range of 0.10 to 0.15 (Gusev).
Such low $F_{ST}$ values indicate that only a small proportion of the total genetic variation in humans is due to differences between subpopulations (10-15%). The remaining, much larger proportion (85-90%) of the genetic variation is found within each subpopulation.
These $F_{ST}$ estimates are one of many strong arguments against the idea of distinct, biologically defined "races." Traditional racial classifications based on superficial characteristics (like skin color) do not reflect underlying patterns of genetic variation. There is far more genetic diversity within so-called racial groups than between them.
TL;DR: SNPs are single nucleotide differences in the genome sequences (Book of Me) of a population. They comprise the bulk of genetic variation in humans. Most SNPs confer no phenotypic changes, but they are valuable genetic markers for genetic association and ancestry studies as described in future posts.
The alleged murderer of UnitedHealth CEO Brian Thompson was recently caught, identified by an observant customer and employee at a Pennsylvania McDonald’s. While he was still on the run, I was confident he would eventually be caught once I heard they had obtained a DNA sample from saliva on a bottle or cup. There has been a recent revolution in DNA forensics, which has led to many cold cases being solved (Wikipedia). In these cases, a sample of DNA from the alleged perpetrator was left at the scene of the crime, but police did not have a person to match the DNA with.
Over the past 20 years, we have developed science fiction level technology to sequence DNA quickly, cheaply, and from a small sample. It is now possible to sequence a whole human genome for close to $100 in a single day. One only needs to remember the sequencing of the first human genome over a period of roughly 10 years (1990-2001) for a cost of approximately $5 billion (present day dollars) to appreciate the insane progress. In the not-too-distant future, we all will have our genomes sequenced.
Intuitively, people understand that the DNA sequence of an individual is unique (excepting identical twins), and that DNA sequences of close relatives will be more similar than the DNA sequences of two strangers, but can we be more precise? For each of the 23 pairs of chromosomes in nearly every cell in the human body, you inherit one chromosome (single molecule of DNA) from Mom (maternal) and one chromosome from Dad (paternal). Is the chromosome you inherit from Mom (say chromosome 1) exactly the same as the chromosome in all of Mom’s cells? The answer is no, because recombination occurs during meiosis (special cell division that produces the gametes) so that the chromosome 1 found in a particular gamete (egg) is a mix of your Mother’s maternal and paternal chromosome 1. During recombination, segments of DNA are exchanged between the two homologous chromosomes in a pair resulting in recombinant chromosomes.
I previously mentioned the concept of identity-by-descent (IBD) which refers to shared, completely identical DNA segments between two or more individuals inherited from a common ancestor. Two siblings share 50% of their DNA via identity-by-descent. Each sibling inherits a chromosome 1 from Mom. Meiotic recombination described in the previous paragraph mixes the chromosome so that half is from Mom’s Mom and half is from Mom’s Dad. This mixing is random so that one sibling may get a particular segment from Mom’s Mom, whereas the other sibling has a 50% chance of getting the identical segment from Mom’s Mom but also a 50% chance of getting the segment from Mom’s Dad. Thus, for siblings, 50% of their maternal chromosome 1 will be shared from a common ancestor (i.e. Mom). The same applies for their paternal chromosome 1.
Without going into detail, for each succeeding generation the IBD sharing decreases by a factor of 4, and so first cousins share $\frac{1}{2} \times \frac{1}{4} = \frac{1}{8} = 12.5\%$ DNA, and 2nd cousins $\frac{1}{8} \times \frac{1}{4} = \frac{1}{32} = 3.125\%$. Briefly, first cousins related through their fathers share 25% of their paternal chromosome 1; their maternal chromosome 1 is 0% IBD since their Moms are not related. Meiotic recombination will produce various first chromosomes that are 50% paternal and 50% maternal which will be passed down to their kids. Any previously shared segment in the paternal chromosome 1 between the cousins has only a 1/2 chance of being passed down in each cousin resulting in IBD sharing that decreases by $\frac{1}{2} \times \frac{1}{2} = \frac{1}{4}$ between the second cousins compared to the first cousins.
Thus, it is straightforward to calculate the amount of DNA shared between various types of relatives as shown in the Table below:
Relationship
Average % DNA Shared
Range (cM)*
Identical Twin
100%
N/A
Parent / Child
50%
~3719
Full Sibling
50%
2826 - 4537
Grandparent / Grandchild
25%
1264 - 2529
Aunt / Uncle
25%
1264 - 2529
Niece / Nephew
25%
1264 - 2529
Half Sibling
25%
1264 - 2529
1st Cousin
12.5%
298 - 1710
2nd Cousin
3.13%
149 - 446
3rd Cousin
0.78%
0 - 164
4th Cousin
0.20%
0 - 60
*1 centimorgan (cM) is approximately 1 million base pairs (1 Mb) in humans.
Beyond 5th cousins, the amount of shared DNA becomes very small and difficult to distinguish from random background sharing. In other words, short identical stretches of DNA may not be inherited from a recent common ancestor, but rather reflect chance similarities.
How can DNA genealogy be used to track down a criminal, i.e. pick out a suspect based on a DNA sample as described in the first paragraph? One can take advantage of large commercial genealogy (ancestry) databases to identify distant relatives of the suspect, e.g 5th cousin or closer. How big of a database do we need, and what is the probability of obtaining a hit?
Yaniv Erlich and colleagues provided approximate answers to these questions in the article "Identity inference of genomic data using long-range familial searches" which describes the use of consumer genomic databases for identifying individuals through distant familial relatives. Combining population genetics theory and computer simulations, the authors reached the following conclusions:
About 60% of searches for individuals of European descent in a database of 1.28 million will find a third-cousin or closer match.
With a database covering just 2% of a target population, it is theoretically possible to find a third-cousin match for nearly anyone in that population.
For U.S. individuals of European descent, a database of approximately 3 million could provide a third-cousin match for over 99% of the population.
Thus, a database of 1 million samples is large enough for this purpose, and many such ancestry databases exist that the public can search for matches (relatives).
What is the next step after a hit has been found in an ancestry database? Using the relative as a starting point, authorities can construct a family tree with the goal of eventually adding the suspect to the tree, i.e. build the tree out to the 3rd (or 4th or 5th) cousins of the hit. Then basic demographic information (e.g. man or woman, age, etc.) can be used to pinpoint the suspect within the tree.
In an investigation, the family tree is often constructed by hand, but in the paper, the authors took advantage of "86 million profiles from publicly-available online data shared by genealogy enthusiasts. After extensive cleaning and validation, we obtained population-scale family trees, including a single pedigree of 13 million individuals."
From these population scale family trees, each hit (in simulations) produced on average a list of ~850 individuals who possessed the appropriate degree of DNA sharing. Then, "localizing the target to within 160 km (100 miles) will exclude 57% of the candidates on average. Next, availability of the target’s age to within ±5 years will exclude 91% of the remaining candidates. Finally, inference of the biological sex of the target will halve the list to just around 16 to 17 individuals, a search space that is small enough for manual inspection.” Thus based on this analysis it should be possible from location, age, and sex, to narrow down the possible suspects to a small handful of people that can be investigated one by one.
In theory, one sequences the DNA sample left behind by an unknown suspect (i.e. hair, saliva, semen, etc.), and uploads the sequence to large commercial ancestry databases containing a million or more entries. One can expect an approximate 3rd cousin hit. A family tree is constructed from this match (either by hand or in an automated fashion taking advantage of online genealogy information) that is likely to encompass the suspect. Demographic information can cull this tree to narrow in on the suspect.
How well does this procedure work in practice? It turns out quite well, and one example is provided by the saga of the Golden Gate Killer, one of the most celebrated cold cases.
The Golden State Killer is a moniker for a serial offender who committed a series of crimes in California between 1974 and 1986, including burglaries, rapes, and murders. Earlier in his criminal career (1974–1975), he was involved in over 100 burglaries in Visalia, often described as the "Visalia Ransacker." In the mid-1970s, the offender was responsible for over 50 sexual assaults in the Sacramento area and other parts of Northern California. Later, in Southern California, he committed at least 13 murders between 1979 and 1986 and was known as the "Original Night Stalker." By 2001, the authorities were able to link the various cases to the same culprit via DNA who became known as the Golden State Killer (GSK). Unfortunately, there were no new leads and so the case went cold.
However, in 2016 investigators made a renewed effort taking advantage of the DNA samples from GSK and advances in DNA forensic genealogy. They uploaded the DNA sequence to several consumer genetic databases (~1-2 million customers each), and obtained 3rd cousin hits. In Feb. 2018, they uploaded the sequence to the MyHeritage website and did even better with a 2nd cousin hit.
Starting with this second cousin, they built a family tree using the consumer genealogy site Ancestry.com. They then obtained DNA from a member of the family tree to exclude certain people in one portion of the tree (i.e. person and his close relatives were not an exact match with GSK DNA). Only six male cousins remained as possible fits. An FBI search of California driver’s license records showed that only one of those six men had blue eyes consistent with the profile (from DNA sequence it is possible to predict blue or brown eyes with reasonable accuracy which will be the topic of a future blog post).
After 10 days of surveillance that included police enlisting the help of a garbage truck driver to snatch DNA-bearing items from his trash can, Joseph James DeAngelo was arrested on April 24, 2018 and charged with multiple counts of murder. They were able to exactly match DeAngelo's DNA to the DNA of the Golden State Killer which provided conclusive proof. He is currently serving 26 life sentences with no possibility of parole. More information on the GSK case can be found in the following LA Times article, and an excellent video by Veritasium (see below).
TL;DR: Authorities can sequence a tiny sample of DNA from a criminal suspect, and then upload that sequence to ancestry databases. For large databases (> 1 million individuals), they can roughly expect a 3rd cousin match or so. A family tree constructed from the match may include the suspect. Basic demographic information can help authorities prune the tree leaving a small number of remaining candidates from which they can obtain DNA to match with the original sample. If the authorities did indeed have a DNA sample from the killer of Brian Thompson, then eventually his DNA would have led to his capture even without the timely aid of an observant public.
There are two types of twins: monozygotic (MZ, identical), and dizygotic (DZ, fraternal). The term monozygotic refers to the one-egg origin of identical twins; during early cleavage, the embryo splits into two separate cell masses each of which gives rise to a twin, both of whom originated from the same fertilized egg (zygote). Dizygotic twins arise from two distinct fertilized eggs, and hence are genetically related as any two siblings. Monozygotic twins are clones; they possess identical genotypes and DNA sequence. Dizygotic twins share 50% DNA identity-by-descent (IBD) typical for siblings.
Fig 1: Monzygotic twins originate from he same fertilized egg (zygote). Dizygotic twins originate from two separate fertilized eggs in which two distinct oocytes were fertilized by different sperm (from Figure 25.7 of "Genetics: From Genes to Genomes", 7th Edition).
For a given trait, one expects MZ twins to be more phenotypically similar than DZ twins depending on the heritability of the trait. If the trait is 100% heritable, then the MZ twins should have identical phenotypes reflecting their identical genotypes, whereas DZ twins would have a phenotypic correlation of 0.5. If the trait is 0% heritable, then the MZ twins should be as similar to each other as any two people that share a home environment such as a pair of DZ twins.
Classical Twin Equation
One can take advantage of the two types of twins to estimate heritabiilty by comparing the phenotypic similarity (covariance) between MZ and DZ twins. But let’s first start with some notation.
Previously we wrote the phenotypic value (the measured value of a trait) $z = \mu + G + E$, in which $\mu$ is the mean value of the trait in the population, $G$ is the genotypic value (contribution of genetic factors) and $E$ is the environmental value (contribution of environmental factors). $G$ and $E$ have been normalized to mean 0, and this equation is one form of the Fundamental Law of Genetics that Phenotype = Genotype + Environment.
$G$ can be broken up into an additive term, dominance term, and epistatic term. Previously we did not decompose $E$, but for twin studies we will divide it into a common environment term (environment that the two twins share) and the unique environment term (environment that they experience separately). This gives us $z = \mu + A + D + S + C + E$ in which $A$ is the additive, $D$ is the dominance, and $S$ the epistatic genotypic terms, and $C$ is the common and $E$ the unique environment terms. As a notational change we will write each of these terms as a standardized random variable with unit variance (upper-case) multiplied by a coefficient (lower-case) so that for example the additive genotypic effect previously denoted as $A$ will now be written as $aA$ with the additive variance equal to $a^2$. Finally, we will be focusing on the simplest of the twin models, the ACE model (Additive-Common-unique Environment), which ignores the dominance and epistatic terms. That leaves us with: $z = aA + cC + eE$.
Just as we calculated the phenotypic covariance between parent and offspring to estimate heritability by parent-offspring regression, we can calculate the covariance between twins ($T_1$ and $T_2$), with $\operatorname{Cov}(T_1,T_2) = \operatorname{Cov}(a_1A_1,a_2A_2) + \operatorname{Cov}(c_1C_1,c_2C_2) + \operatorname{Cov}(e_1E_1,e_2E_2) + \text{cross terms}$. For now we assume the cross-terms e.g. $\operatorname{Cov}(a_1A_1,c_2C_2)$ are 0, which we will discuss below. We also assume that the unique environment covariance is 0: $\operatorname{Cov}(e_1E_1,e_2E_2) = 0$ (e.g. the unique friends of the two twins are uncorrelated). The common environment for twins is the same so that $\operatorname{Cov}(c_1C_1,c_2C_2) = \operatorname{Var}(cC) = c^2$. This leaves us with the additive genotypic covariance which for identical twins should be $\operatorname{Cov}_{MZ}(a_1A_1,a_2A_2) = \operatorname{Var}(aA) = a^2$ because they possess the same genotype. Dizygotic twins are like any two siblings who share 50% of their genotype identity-by-descent (IBD) as do parent and offspring, and so $\operatorname{Cov}_{DZ}(a_1A_1,a_2A_2) = \frac{a^2}{2}$ using the same reasoning as when we calculated parent-offspring covariance. Thus we have: $ \begin{equation} \operatorname{Cov}_{MZ}(T_1,T_2) = a^2 + c^2 \text{ and } \operatorname{Cov}_{DZ}(T_1,T_2) = \frac{a^2}{2} + c^2 \end{equation} $. The phenotypic variance for any twin is the same as the variance for any member of the population which is $V_p = a^2 + c^2 + e^2$ given the assumptions of the ACE model. We can convert the covariances to correlations ($r$) by dividing by the variance:
If we subtract the two correlations, then the common environment term cancels out, and we are left with the additive genotypic value divided by 2. Multiplying by 2 results in the following expression which is equal to narrow-sense heritability:
Thus, 2 times the difference between the monozygotic twin trait correlation and the dizygotic twin trait correlation equals narrow-sense heritability. This is the classical twin equation which is sometimes referred to as Falconer's equation.
Performing a Classical Twin Study
Operationally, performing a twin study is straightforward. One collects data from a set of MZ twins and from a set of DZ twins by measuring the trait of interest in all twins. Then one calculates the correlation between pairs of twins by plotting the trait value of twin 1 on the x-axis and the trait value of twin 2 on the y-axis followed by fitting a line to the points. The slope of the regression line $b = r \frac{\sigma_2}{\sigma_1}$ is equal to the correlation coefficient $r$ multiplied by the ratio of the standard deviation of twin 2 trait values ($\sigma_2$) over the standard deviation for twin 1 trait values ($\sigma_1$). Thus, we can obtain $r_{MZ}$ and $r_{DZ}$ by plotting the two sets of twins on separate graphs in an Excel spreadsheet. Plug the resulting values into the twin equation (Eq 2), and voila you have a heritability estimate for the trait and potentially a paper. This simplicity has resulted in twin studies becoming a cottage industry in the human genetics field.
ACE Model is an Oversimplification
But life is not so simple, or as Sasha Gusev so eloquently puts it:
Notice all of the extra terms compared to Eqs 1a and 1b of the ACE model. There is the dominance term $D$, an epistatic term ($A*A$), a term for gene-environment interaction ($A*C$), a term that takes into account assortative mating and gene-environment correlation ($r_A$), and separate terms for the common environments of MZ versus DZ twins ($C_{MZ}$ and $C_{DZ}$). Subtracting the two correlation coefficients now results in a much more complicated expression than $\frac{a}{2}$ with extra terms that can potentially inflate the value of $2(r_{MZ} - r_{DZ})$.
It is important to emphasize that the common environment term for MZ twins ($C_{MZ}$) will not be the same as the common environment term for DZ twins ($C_{DZ}$) despite being written as the same $c$ in Eq 1. A priori, one expects the environment for MZ twins to be more similar because they are more phenotypically similar in numerous traits other than the one whose heritability is being estimated. This greater phenotypic similarity should attract and create greater environmental similarity. For example, the heritability of height is relatively large compared to other traits, and so MZ twins are more likely to be the same height, and as a result both may play basketball which can produce more common environment effects compared to DZ twins who are of different heights so that one plays basketball and the other does not. Thus, $C_{MZ}$ should be bigger than $C_{DZ}$ leading to an artifactually inflated value for $r_{MZ}$, and hence heritability. In most twin studies including the ACE model, the assumption is $C_{MZ} = C_{DZ}$ so that it can be written just as $C$ (or $c$) which cancels out in the numerator of Eq 2.
The latter is commonly referred to as the Equal Environments Assumption (EEA), which posits that monozygotic (identical) twins and dizygotic (fraternal) twins share their common environments to the same extent. If monozygotic twins are treated more similarly or share more environmental influences than dizygotic twins, this could lead to an overestimation of heritability.
A second key assumption of the classical twin study is completely ignoring gene-environment interaction and correlation (some of the cross-terms in covariance calculation). If there is gene-environment interaction (in which the effect of genes on a trait depends on the environment, or vice versa; the $A*C$ term), or gene-environment correlation (in which genetic factors influence exposure to environmental factors), this could also upwardly bias heritability estimates. For example, if gene-environment correlations exist, then more similar genotypes can lead to more similar environments even for the "unique" environmental factors.
Finally, it is possible that non-equal environments and gene-environment interactions can feed back upon each to further increase the difference between $r_{MZ}$ and $r_{DZ}$ resulting in even more upward bias. One DZ twin ends up on the basketball team whereas the other ends up on the chess team, and this difference reinforces non-equal environments because of the different circles of friends.
TL;DR: The classical twin study is straightforward to describe and perform but depends on numerous questionable assumptions that upwardly bias the heritability estimates, including EEA and neglecting to consider gene-environment effects.
The fundamental problem is that $a$, $c$, and $e$ in Eq 1 are latent variables that cannot be experimentally measured directly and so instead are inferred indirectly from a model (e.g. ACE model), but the model possesses suspect assumptions so that the inferred values are also suspect.
In contrast, one may adopt the more direct approach of predicting phenotype from genotype, with the accuracy of this prediction being limited by the heritability of the trait. The more heritable the trait, the more accurate the prediction.
Given that we know the complete genotype of an individual from the DNA sequence, then we should be able to make a phenotypic prediction for any trait as long as we have collected enough data to construct a powerful enough (linear) model relating genotype to phenotype. This idea forms the basis of molecular heritability estimates which is the subject of a future post and whose values tend to be significantly lower than those from classical twin studies.