The chapters will come in pairs representing the homologous chromosomes, e.g. maternal and paternal chromosome 1, which would represent the closely related Chapter 1M and Chapter 1P. There are approximately 20,000 paragraphs spread across each set of 23 (maternal or paternal) chapters (or ~40,000 paragraphs in total). Like the chapters, the paragraphs come in pairs with a particular paragraph in Chapter 1M, being nearly identical to the corresponding paragraph in Chapter 1P.
You can compare your Book of Me with that of your friend or sibling. In a previous post, I mentioned that you share 50% of your DNA (identity-by-descent, IBD) with a full sibling. Does that mean that 50% of your sequence will be identical to the corresponding sequence in your brother or sister? No.
Instead, you find that your sequence is approximately 99.95% identical to that of your sibling (excluding differences between X and Y sex chromosomes). Perhaps, even more surprisingly, the sequence in your Book of Me will be 99.9% identical with that of a random stranger, no matter what the race or ethnicity. Between any two random individuals there is roughly one character difference every 1000 characters (base pairs; because DNA is double-stranded we speak of a single position on DNA as comprising a base pair, which are the two nucleotides on the complementary strands hydrogen-bonded to each other).
![]() |
| A SNP is a single nucleotide (character) difference between the sequences of two individuals. Note we show the sequence from one of the two DNA strands in the 5' to 3' direction. |
This sequence difference is referred to as a single nucleotide variant (SNV) or alternatively a single nucleotide polymorphism (SNP), which is defined as a variation at a single position in a DNA sequence among individuals of a species. Technically speaking, a SNP occurs when a single nucleotide (A, T, C, or G) is altered at a specific location in the genome that is found in at least 1% of the population, i.e. it is relatively common. A SNV refers to any single nucleotide change, regardless of its frequency, i.e. it can be rare. I will use the terms SNP and SNV interchangeably.
Aside from SNPs, there will be other types of sequence differences between two individuals representing other sources of genetic variation. At the largest scale, there are chromosome level changes such as chromosome deletions, inversions, duplications, etc. that can span millions of base pairs or even a whole chromosome (aneuplody = having extra or missing chromosomes such as trisomy 21 which is Down Syndrome). At the smallest scale is the aforementioned SNP at the single base pair level.
In between, are the InDels, which are short (1-100bp) insertions or deletions, and the simple sequence repeats which are short (1-10bp) sequences repeated up to 100 times. There are also larger (thousands of base pairs) DNA segments that can encompass whole genes that are repeated called Copy Number Variants. Of these different classes of genetic variants, SNPs are by far the most prevalent at roughly 1 every 1000 base pairs. The others are 10 to 100 fold less frequent. Thus, when we speak of genetic variation we are mainly talking about SNPs.
SNPs are typically biallelic (two alleles), a major allele (more common) and a minor allele (less common). In total there are 4 possible alleles (the 4 possible nucleotides), but not enough time has passed since the first modern humans (roughly 100,000 years) for random drift to sample all possible nucleotides at each position. The polymorphism is usually reported as the minor allele frequency (MAF). The smaller the MAF, the less polymorphic, i.e. most people in the population will possess the major allele and thus be identical at the locus.
One can consider the minor allele to be a mutant allele in a certain sense, and ask whether this mutant allele has a corresponding mutant phenotype. The larger mutations, i.e. chromosomal rearrangements, are expected to have a significant phenotypic impact because they can alter the expression of one or more genes. A SNP, on the other hand, perturbs only a single base pair, and typically the minor allele does not produce a phenotype different from the wild-type (major allele) base pair. The principal reason is that the SNP locus (position) is typically in the intergenic region (between genes) or in introns, neither of which may have a functional role. Indeed the protein coding region of genes represents only 1.5 - 2% of the human genome. Even if located within the gene coding region, the SNP may not alter the protein sequence (e.g. silent mutation), or may make a conservative amino acid substitution that has minimal effect on protein regulation or function.
SNPs play a central role in human genetics because they can serve as genetic markers. Although any given SNP (technically speaking, SNP allele) may not be the causal mutation of some human trait of interest, SNPs can be in close proximity (linked) to the causal mutation, which could be a different SNP or combination of SNPs or one of the other types of genetic variants (e.g. InDel). Because of this linkage, the SNP will associate with the mutant phenotype in a population, and segregate with the phenotype in a pedigree (recombination is unlikely to separate a particular SNP allele from the causal mutation). Genotyping (i.e. sequencing) the SNP to determine the allele (major or minor) in an individual can thus indicate the presence of the linked causal mutation and phenotype.
As the main source of genetic variation, SNPs are also critical for characterizing human ancestry and diversity. The first sequenced human genome ultimately became known as the reference genome. Subsequent genome sequences could be compared to this reference sequence as well as to each other. As mentioned above, any two genomes differ at approximately 1 in 1000 base pairs, but more related individuals are expected to have fewer differences than less related. Three of the earliest sequenced genomes were those of James Watson, Craig Venter, and an Asian man with initials YK (also involved in the Human Genome Project). Each of the genomes were compared to the reference genome and to each other with the figure below showing the number of SNP differences compared to the reference genome.
One surprise was that the number of SNPs shared between Venter and Watson, two Caucasians, was 1.7 million, whereas Venter and YH shared 1.6M, and Watson and YH shared 1.65 million. A naive expectation was that Watson and Venter would share far more SNPs than they shared with the Asian YH. Instead the differences appeared to be small (50,000-100,000) especially compared to the large number of SNPs shared between YH and the other two, roughly 1.6 million each.
This observation hints that the so-called races may not be so different at the genetic level. We can formalize the insight using Wright’s fixation index ($F_{ST}$), which is a measure of genetic differentiation between subpopulations relative to the total population. According to Gusev, the classic definition of $F_{ST}$ for a single site is "the correlation between gametes chosen randomly from within the same subpopulation relative to the [total] population." It quantifies the proportion of genetic variance that is due to differences between subpopulations compared to within subpopulations.
Thanks to the automated DNA sequencing revolution, hundreds of thousands of human genomes have now been sequenced. As a result, researchers have been able to calculate $F_{ST}$ for various pairs of subpopulations. Among geographically defined groups (e.g. East Asians, sub-Saharan Africans, Europeans), $F_{ST}$ between any two was estimated to be in the range of 0.10 to 0.15 (Gusev).
Such low $F_{ST}$ values indicate that only a small proportion of the total genetic variation in humans is due to differences between subpopulations (10-15%). The remaining, much larger proportion (85-90%) of the genetic variation is found within each subpopulation.
These $F_{ST}$ estimates are one of many strong arguments against the idea of distinct, biologically defined "races." Traditional racial classifications based on superficial characteristics (like skin color) do not reflect underlying patterns of genetic variation. There is far more genetic diversity within so-called racial groups than between them.
TL;DR: SNPs are single nucleotide differences in the genome sequences (Book of Me) of a population. They comprise the bulk of genetic variation in humans. Most SNPs confer no phenotypic changes, but they are valuable genetic markers for genetic association and ancestry studies as described in future posts.


