Friday, January 31, 2025

DNA forensic genealogy solved the Golden Gate Killer case

The alleged murderer of UnitedHealth CEO Brian Thompson was recently caught, identified by an observant customer and employee at a Pennsylvania McDonald’s. While he was still on the run, I was confident he would eventually be caught once I heard they had obtained a DNA sample from saliva on a bottle or cup. There has been a recent revolution in DNA forensics, which has led to many cold cases being solved (Wikipedia). In these cases, a sample of DNA from the alleged perpetrator was left at the scene of the crime, but police did not have a person to match the DNA with.

Over the past 20 years, we have developed science fiction level technology to sequence DNA quickly, cheaply, and from a small sample. It is now possible to sequence a whole human genome for close to $100 in a single day. One only needs to remember the sequencing of the first human genome over a period of roughly 10 years (1990-2001) for a cost of approximately $5 billion (present day dollars) to appreciate the insane progress. In the not-too-distant future, we all will have our genomes sequenced.

Intuitively, people understand that the DNA sequence of an individual is unique (excepting identical twins), and that DNA sequences of close relatives will be more similar than the DNA sequences of two strangers, but can we be more precise? For each of the 23 pairs of chromosomes in nearly every cell in the human body, you inherit one chromosome (single molecule of DNA) from Mom (maternal) and one chromosome from Dad (paternal). Is the chromosome you inherit from Mom (say chromosome 1) exactly the same as the chromosome in all of Mom’s cells? The answer is no, because recombination occurs during meiosis (special cell division that produces the gametes) so that the chromosome 1 found in a particular gamete (egg) is a mix of your Mother’s maternal and paternal chromosome 1. During recombination, segments of DNA are exchanged between the two homologous chromosomes in a pair resulting in recombinant chromosomes.

I previously mentioned the concept of identity-by-descent (IBD) which refers to shared, completely identical DNA segments between two or more individuals inherited from a common ancestor. Two siblings share 50% of their DNA via identity-by-descent. Each sibling inherits a chromosome 1 from Mom. Meiotic recombination described in the previous paragraph mixes the chromosome so that half is from Mom’s Mom and half is from Mom’s Dad. This mixing is random so that one sibling may get a particular segment from Mom’s Mom, whereas the other sibling has a 50% chance of getting the identical segment from Mom’s Mom but also a 50% chance of getting the segment from Mom’s Dad. Thus, for siblings, 50% of their maternal chromosome 1 will be shared from a common ancestor (i.e. Mom). The same applies for their paternal chromosome 1.

Without going into detail, for each succeeding generation the IBD sharing decreases by a factor of 4, and so first cousins share $\frac{1}{2} \times \frac{1}{4} = \frac{1}{8} = 12.5\%$ DNA, and 2nd cousins $\frac{1}{8} \times \frac{1}{4} = \frac{1}{32} = 3.125\%$. Briefly, first cousins related through their fathers share 25% of their paternal chromosome 1; their maternal chromosome 1 is 0% IBD since their Moms are not related. Meiotic recombination will produce various first chromosomes that are 50% paternal and 50% maternal which will be passed down to their kids. Any previously shared segment in the paternal chromosome 1 between the cousins has only a 1/2 chance of being passed down in each cousin resulting in IBD sharing that decreases by $\frac{1}{2} \times \frac{1}{2} = \frac{1}{4}$ between the second cousins compared to the first cousins. 

Thus, it is straightforward to calculate the amount of DNA shared between various types of relatives as shown in the Table below:

Relationship     Average % DNA Shared      Range (cM)*
Identical Twin 100% N/A
Parent / Child 50% ~3719
Full Sibling 50% 2826 - 4537
Grandparent / Grandchild 25% 1264 - 2529
Aunt / Uncle 25% 1264 - 2529
Niece / Nephew 25% 1264 - 2529
Half Sibling 25% 1264 - 2529
1st Cousin 12.5% 298 - 1710
2nd Cousin 3.13% 149 - 446
3rd Cousin 0.78% 0 - 164
4th Cousin 0.20% 0 - 60
*1 centimorgan (cM) is approximately 1 million base pairs (1 Mb) in humans.

Beyond 5th cousins, the amount of shared DNA becomes very small and difficult to distinguish from random background sharing. In other words, short identical stretches of DNA may not be inherited from a recent common ancestor, but rather reflect chance similarities.

How can DNA genealogy be used to track down a criminal, i.e. pick out a suspect based on a DNA sample as described in the first paragraph? One can take advantage of large commercial genealogy (ancestry) databases to identify distant relatives of the suspect, e.g 5th cousin or closer. How big of a database do we need, and what is the probability of obtaining a hit?

Yaniv Erlich and colleagues provided approximate answers to these questions in the article "Identity inference of genomic data using long-range familial searches" which describes the use of consumer genomic databases for identifying individuals through distant familial relatives. Combining population genetics theory and computer simulations, the authors reached the following conclusions:
  • About 60% of searches for individuals of European descent in a database of 1.28 million will find a third-cousin or closer match.
  • With a database covering just 2% of a target population, it is theoretically possible to find a third-cousin match for nearly anyone in that population.
  • For U.S. individuals of European descent, a database of approximately 3 million could provide a third-cousin match for over 99% of the population.
Thus, a database of 1 million samples is large enough for this purpose, and many such ancestry databases exist that the public can search for matches (relatives).

What is the next step after a hit has been found in an ancestry database? Using the relative as a starting point, authorities can construct a family tree with the goal of eventually adding the suspect to the tree, i.e. build the tree out to the 3rd (or 4th or 5th) cousins of the hit. Then basic demographic information (e.g. man or woman, age, etc.) can be used to pinpoint the suspect within the tree. 

In an investigation, the family tree is often constructed by hand, but in the paper, the authors took advantage of "86 million profiles from publicly-available online data shared by genealogy enthusiasts. After extensive cleaning and validation, we obtained population-scale family trees, including a single pedigree of 13 million individuals."

From these population scale family trees, each hit (in simulations) produced on average a list of ~850 individuals who possessed the appropriate degree of DNA sharing. Then, "localizing the target to within 160 km (100 miles) will exclude 57% of the candidates on average. Next, availability of the target’s age to within ±5 years will exclude 91% of the remaining candidates. Finally, inference of the biological sex of the target will halve the list to just around 16 to 17 individuals, a search space that is small enough for manual inspection.” Thus based on this analysis it should be possible from location, age, and sex, to narrow down the possible suspects to a small handful of people that can be investigated one by one.

In theory, one sequences the DNA sample left behind by an unknown suspect (i.e. hair, saliva, semen, etc.), and uploads the sequence to large commercial ancestry databases containing a million or more entries. One can expect an approximate 3rd cousin hit. A family tree is constructed from this match (either by hand or in an automated fashion taking advantage of online genealogy information) that is likely to encompass the suspect. Demographic information can cull this tree to narrow in on the suspect.

How well does this procedure work in practice? It turns out quite well, and one example is provided by the saga of the Golden Gate Killer, one of the most celebrated cold cases. 

The Golden State Killer is a moniker for a serial offender who committed a series of crimes in California between 1974 and 1986, including burglaries, rapes, and murders. Earlier in his criminal career (1974–1975), he was involved in over 100 burglaries in Visalia, often described as the "Visalia Ransacker." In the mid-1970s, the offender was responsible for over 50 sexual assaults in the Sacramento area and other parts of Northern California. Later, in Southern California, he committed at least 13 murders between 1979 and 1986 and was known as the "Original Night Stalker." By 2001, the authorities were able to link the various cases to the same culprit via DNA who became known as the Golden State Killer (GSK). Unfortunately, there were no new leads and so the case went cold.

However, in 2016 investigators made a renewed effort taking advantage of the DNA samples from GSK and advances in DNA forensic genealogy. They uploaded the DNA sequence to several consumer genetic databases (~1-2 million customers each), and obtained 3rd cousin hits. In Feb. 2018, they uploaded the sequence to the MyHeritage website and did even better with a 2nd cousin hit.

Starting with this second cousin, they built a family tree using the consumer genealogy site Ancestry.com. They then obtained DNA from a member of the family tree to exclude certain people in one portion of the tree (i.e. person and his close relatives were not an exact match with GSK DNA). Only six male cousins remained as possible fits. An FBI search of California driver’s license records showed that only one of those six men had blue eyes consistent with the profile (from DNA sequence it is possible to predict blue or brown eyes with reasonable accuracy which will be the topic of a future blog post).

After 10 days of surveillance that included police enlisting the help of a garbage truck driver to snatch DNA-bearing items from his trash can, Joseph James DeAngelo was arrested on April 24, 2018 and charged with multiple counts of murder. They were able to exactly match DeAngelo's DNA to the DNA of the Golden State Killer which provided conclusive proof. He is currently serving 26 life sentences with no possibility of parole. More information on the GSK case can be found in the following LA Times article, and an excellent video by Veritasium (see below).

TL;DR: Authorities can sequence a tiny sample of DNA from a criminal suspect, and then upload that sequence to ancestry databases. For large databases (> 1 million individuals), they can roughly expect a 3rd cousin match or so. A family tree constructed from the match may include the suspect. Basic demographic information can help authorities prune the tree leaving a small number of remaining candidates from which they can obtain DNA to match with the original sample. If the authorities did indeed have a DNA sample from the killer of Brian Thompson, then eventually his DNA would have led to his capture even without the timely aid of an observant public.


ChatGPT identified 142 statistical "issues" in problematic paper

Today the Bayesian statistician Andrew Gelman posted a rather pointed critique of a paper on his blog stating: "Wow! This paper is an ...