Genetic variation and molecular markers

Two full siblings receive chromosomes from the same parents, yet their DNA sequences are usually not identical. The meiosis explains how recombination reshuffles parental chromosomes. Recombination changes which existing alleles travel together, but it does not create those alleles. To understand why individuals carry different sequence versions in the first place, we need to examine the origin and forms of genomic variation.

Mutation creates new sequence variants

A mutation is a change in a DNA sequence. It may arise through an error during DNA replication, damage to DNA, or imperfect repair. Once a mutation has occurred, later DNA replication can copy the changed sequence. Mutation is therefore the original source of new sequence variants, whereas recombination and segregation redistribute variants that already exist.

From a base substitution to a population SNP

The words mutation, substitution, single-nucleotide variant, and SNP answer different questions. A mutation is an event or process that changes a DNA sequence in one cell lineage. A base substitution is one possible kind of mutation: one base is replaced by another at a particular genomic position. For example, if a germline cell changes from A to G at one position, that event is a substitution mutation. It can be copied into a gamete and inherited by descendants, but it begins as one event, not as a population-wide property.

A substitution between two purines (A and G) or between two pyrimidines (C and T) is a transition. A substitution between a purine and a pyrimidine is a transversion. These labels describe the chemical classes of the bases that changed. They do not say whether the change is common, ancestral, harmful, favourable, or causal for a trait.

A single-nucleotide variant (SNV) is the observed sequence difference at one base position when one sequence, individual, or allele is compared with another. The A/G difference above is an SNV whether it is observed once or many times. The word SNP, short for single-nucleotide polymorphism, adds a population perspective: it refers to a single-base site at which more than one allele occurs among individuals in a population. In some human genetics resources, an allele frequency of at least 1% is used as a practical SNP convention. That threshold is not a universal biological boundary and must not be transferred uncritically to a breeding population with its own sample size, ancestry, and purpose.

The distinction can be summarised as a sequence of questions. Did a DNA change occur in a lineage? That is a mutation. Was the change one base replaced by another? That is a substitution. Does a comparison reveal a one-base difference? That is an SNV. Do two or more alleles occur at that one-base site in the population being described? It is a polymorphic site and may be called an SNP under the stated convention. One substitution can create an allele later observed as an SNP, but the terms are not synonyms.

Six individuals each have two chromosome-2 copies. A yellow vertical band marks one sequence position, where the displayed alleles are T, C, or A across the chromosome copies.
Figure 1: A sequence illustration compares the two chromosome copies of six individuals at one highlighted locus. The alleles are T, C, and A; some individuals carry two matching copies and others carry different copies.

Where the change occurs determines whether it can contribute to inherited variation. A germline variant occurs in a cell lineage that contributes to gametes and can be transmitted to offspring. A somatic variant occurs in another cell lineage. In the usual animal-breeding setting, a somatic variant can affect the descendant cells within that individual but is not transmitted through its sperm or eggs. The boundary between reproductive and somatic lineages depends on the organism and mode of reproduction, so this distinction must be interpreted carefully in plants and clonally propagated populations.

An inherited variant was already present in a parent and passed through a gamete. A de novo variant is newly observed in an offspring rather than in the sampled parental genotypes. It may have arisen while a parental germ cell was forming or early in the offspring’s development. Calling a variant de novo also depends on measurement: low coverage, genotyping error, or parental mosaicism can make an inherited change appear absent from the parental samples.

Variant classes describe the scale and form of change

The DNA section introduced a sequence as an ordered series of bases. Variants can alter one base, a few adjacent bases, the number of copies of a region, or the order and location of much larger segments. The schematic examples below use letters and blocks only to show the form of each change.

Variant class Conceptual comparison What changed
Single-nucleotide variant (SNV), including an SNP when the population criterion is met ...A... becomes ...G... One base differs at one genomic position; a substitution is one mechanism that can create this difference.
Insertion or deletion (indel) ACGT becomes ACGGT, or ACGT becomes ACT One or more adjacent bases are added or removed.
Copy-number variant (CNV) A-B-C becomes A-B-B-B-C The number of copies of a DNA segment differs.
Structural variant A-B-C-D becomes A-C-B-D, or a block moves elsewhere A larger segment is deleted, duplicated, inverted, inserted, or translocated.

These classes describe different aspects of sequence change and can overlap. A deletion is an indel when it is short under the convention used by an analysis, but a sufficiently large deletion is commonly treated as a structural variant. A CNV is also a form of structural variation. Size thresholds and reporting conventions can differ among technologies and software, so the operational definition should accompany an analysis.

Variant size does not determine biological importance. A one-base change can disrupt a coding or regulatory sequence, while a large change may have little measured effect in a particular environment. Conversely, most observed sequence variants should not be assumed to change a trait. The effect depends on location, molecular consequence, genetic background, and the phenotype being studied.

Variant, polymorphism, causal variant, and marker are different roles

A variant is a sequence difference relative to another sequence or a chosen reference. An SNV is the single-base member of this broad class. Polymorphism emphasises that two or more sequence forms occur among individuals or populations. An SNP is therefore a polymorphic single-base site, whereas a single-base substitution observed in one individual may still be a rare SNV. Some fields have historically used frequency thresholds to separate a polymorphism from a rare variant, but no single threshold is appropriate for every population and purpose. This curriculum therefore uses variant as the general term and states frequency explicitly when it matters.

A causal variant contributes to a biological process that changes the trait under study. Establishing that role requires evidence beyond statistical association. A nearby variant may show a similar population pattern because the two variants tend to be inherited together, even if only one participates in the mechanism.

A molecular marker is an identifiable DNA feature at a known genomic location whose inheritance can be followed. A marker can lie within a gene, outside a gene, or have no known function. Its value in genomic prediction comes from the information it carries about inherited chromosome segments. A marker may itself be causal, may track a nearby causal variant, or may carry little useful information for a particular trait and population. Marker and causal variant are therefore roles supported by evidence, not interchangeable names.

Reference and alternative alleles are coding labels

A reference sequence provides coordinates and a comparison sequence. At a biallelic site, the allele matching that sequence is often labelled reference, and the other observed allele is labelled alternative. Suppose the reference base at a site is A and a breeding candidate carries A/G. The label tells us which base matches the chosen reference and allows the genotype to be encoded. It does not tell us that A is ancestral, common, favourable, or biologically inactive. It also does not tell us that G is derived, rare, harmful, or causal.

Those biological and population properties require separate evidence. The alternative allele can be the common allele in a particular population, and changing the reference assembly or reversing the counted allele changes the code without changing the individual’s DNA. Later pages will define allele frequency and numerical dosage. For now, the important point is that coding conventions must not be mistaken for biological conclusions.

The usual SNP dosage is especially simple because one biallelic locus has two allele states, so a diploid genotype can count 0, 1, or 2 copies of a chosen allele. A microhaplotype is different: it combines the phased states of several nearby variants on one chromosome copy. The resulting local marker can have more than two observed haplotype alleles, but it is not a third allele at any one constituent SNP. The later page on haplotypes, phasing, and microhaplotypes explains how that multiallelic representation arises and why its usefulness depends on the population, phase quality, and local genomic context.

Sequencing and genotyping measure variation differently

DNA sequencing determines the order of bases across the DNA regions that an assay covers. Those reads can reveal variants relative to a reference sequence, subject to coverage, alignment, and variant-calling uncertainty. Sequencing may target one region, a set of genes, or much of a genome; the word does not by itself guarantee that every base was measured accurately.

Genotyping determines which allele or genotype an individual carries at specified loci. A marker array, for example, can measure a predefined set of SNPs efficiently across many candidates, but it does not directly inspect every unassayed position between them. Genotypes can also be called from sequence data, so sequencing and genotyping are not mutually exclusive labels for instruments. The useful distinction is between reading sequence across assayed regions and assigning genotypes at defined variant sites.

These measurements require quality control before modelling. File layouts, missing-value codes, dosage matrices, and software-ready SNP inputs belong in the SNP input guide, rather than on this biological page.

Observed markers and causal truth in the project data

The recurring project dataset is simulated. Its SNP map records 500 observed markers with identifiers, chromosomes, and base-pair positions. The simulation also records which of those markers generated the genetic signal. This second piece of information is available because the data-generating process is known, not because an ordinary genotyping assay reveals causality.

Code
d <- readRDS("../../demo-data/main/demo_data.rds")
marker_map <- transform(
  d$map_snp,
  causal_in_simulation = seq_len(nrow(d$map_snp)) %in% d$qtl$snp_idx
)

c(
  observed_markers = nrow(marker_map),
  causal_in_simulation = sum(marker_map$causal_in_simulation)
)
    observed_markers causal_in_simulation 
                 500                   20 
Code
table(marker_map$causal_in_simulation)

FALSE  TRUE 
  480    20 
Code
marker_map[
  marker_map$causal_in_simulation,
  c("SNP", "CHROM", "POS")
]
       SNP CHROM      POS
29  SNP029     1  3800000
30  SNP030     1  3900000
74  SNP074     1  8300000
90  SNP090     1  9900000
158 SNP158     2  6700000
188 SNP188     2  9700000
192 SNP192     2 10100000
201 SNP201     3  1000000
242 SNP242     3  5100000
266 SNP266     3  7500000
279 SNP279     3  8800000
282 SNP282     3  9100000
296 SNP296     3 10500000
318 SNP318     4  2700000
334 SNP334     4  4300000
356 SNP356     4  6500000
444 SNP444     5  5300000
445 SNP445     5  5400000
460 SNP460     5  6900000
479 SNP479     5  8800000

The output separates 20 simulated causal SNPs from 480 other observed SNPs. The latter group should not be called biologically irrelevant: a non-causal marker may still predict a causal chromosome segment, and the usefulness of that association can depend on the population. Conversely, inclusion among the 20 causal indices is simulation truth for this teaching dataset, not a claim that real breeding data contain exactly 20 causal loci.

The map contains marker identifiers and positions, but it does not contain the nucleotide identities used as reference and alternative alleles. We can therefore discuss marker locations and the simulated causal set, but we cannot infer an A/G coding or label a favourable allele from this object. Doing so would add information that the dataset does not provide.

Exercises

  1. A sequence comparison changes ACCT to ACGT at one position. Classify the change as a substitution and an SNV. Under what additional population-level condition could the site be called an SNP? Does any of those labels show that the changed base causes a trait difference?

The comparison contains a C-to-G substitution, which is a transversion, and it is an SNV because one base differs. Calling it an SNP requires evidence that more than one allele occurs at that site in the relevant population under the stated convention. None of these labels provides evidence by itself that the site affects a trait.

  1. At one locus, C matches the reference sequence and T is labelled the alternative allele. List three conclusions that cannot be obtained from those labels alone.

The labels alone do not identify which allele is ancestral, which is more common in the target population, or which has a favourable trait effect. They also do not establish whether either allele is causal. Each conclusion needs population, evolutionary, functional, or trait evidence beyond the reference comparison.

  1. The project map contains 500 observed SNP markers, of which 20 are marked causal by the simulation. How many observed markers are not causal in the generating model, and why might some of them still contribute to genomic prediction?

There are \(500-20=480\) observed markers outside the simulated causal set. Some may carry information about nearby causal variants when their alleles are inherited together in the population. The next pages develop the population frequencies and multi-locus associations needed to explain when that marker information persists.

  1. A DNA change is found in a skin cell but not in the individual’s gametes. Is it usually inherited by offspring? Contrast mutation with recombination as a source of genetic variation.

This is a somatic variant, so in the usual animal-breeding setting it is not transmitted through sperm or eggs. Mutation changes a DNA sequence and is the source of new variants. Recombination rearranges variants already present on parental chromosomes into new combinations.

  1. A marker array reports genotypes at predefined SNPs for many candidates. What information can it provide, and what does it not directly determine about unassayed sequence positions?

The array can determine which allele or genotype each candidate carries at its specified marker loci. It does not directly read every sequence position between those markers. Sequencing reads bases across the regions that an assay covers, although both measurements still require quality control.

Identifying a variant in one candidate does not yet tell us whether it is rare, common, or fixed in the breeding population. Those descriptions require counts across individuals. Genetic variation in populations is therefore the next step from molecular differences to the population-level patterns used in genomic prediction.