Overview of Selective Breeding

Muscular cow

Selective breeding has reshaped production animals over generations.

Modern dairy cows produce many times more milk than their ancestors did a century ago. Broiler chickens grow faster, and farmed salmon reach market size sooner. We might assume that these improvements come from the genetically modified organisms (GMOs) that people worry about. Yet, these advancements are actually the result of selective breeding.

Selective breeding aims to produce offspring with specific desired traits by choosing parents that already carry those traits. It can therefore only work with the traits and genetic variation that already exist in a population. This is why diversity matters. If a population holds no variation for a trait, there is nothing to select on.

Challenges in traditional selection

Traditional selection chooses parents by looking only at their own visible traits. This works very well for simple traits, but it struggles with several important ones. Some traits appear late in life, and we must wait years before we can measure them. Some can only be measured after an animal is slaughtered, such as meat quality. Others are expressed in only one sex, such as milk yield in females.

Disease resistance is a good example. To find individuals that tolerate a certain disease, we have to expose them to it, and only the survivors are kept as candidate parents. This is wasteful, slow, and destructive. We may lose valuable individuals to the test itself, and the benefit of selection can take a very long time to arrive.

Advances in molecular biology enable genomic selection

New molecular technologies, particularly genomic sequencing, can now read an individual’s genome sequence. This has shifted breeding from traditional selection toward genomic selection. The statistical framework behind that shift was set out by Meuwissen, Hayes and Goddard (2001), who showed that genetic merit can be predicted from genome-wide dense marker maps, and almost every genomic prediction method in use today still builds on that paper.

The idea rests on a simple assumption. Certain genetic variants in the genome are linked to certain traits. These variants are called quantitative trait loci (QTLs), and they are passed from parents to offspring.

In practice, however, we never know exactly where the QTLs are. They may be many, and they may be scattered across the genome, so genotyping them directly is almost impossible.

Instead, we use genetic markers to point to QTLs. The most common markers are single-nucleotide polymorphisms (SNPs). A marker works because it sits physically close to a QTL on the chromosome. Because they are neighbours, the marker and the QTL tend to be inherited together as a package. This tendency is called linkage disequilibrium (LD).

Single Nucleotide Polymorphisms

A marker is like a signpost on a road. The signpost is not the destination itself, but because it always stands right next to the destination, it tells us the destination is there. In the same way, a marker is not the QTL, but it tells us a QTL is nearby.

Marker and QTL in linkage disequilibrium

A marker sits close enough to a QTL that the two are inherited together.

So instead of measuring the trait itself, we can genotype an individual’s markers and use them to predict how that individual will perform.

How does genomic selection actually work?

Let’s start with a simple case: a single marker.

Consider two individuals. One carries the allele ATC, and the other carries AGC. The individual with ATC has very low resistance to a disease. The individual with AGC has very high resistance, and this pattern is consistent across the whole population.

From this, we can assume that any individual carrying AGC will resist the disease better than one that does not. So we can select AGC carriers as parents and pass that allele on to their offspring.

In reality, it is rare for a single allele to control a trait. Most production traits, such as disease resistance, growth rate, and meat quality, are driven by many QTLs at once. This means many markers are involved, and each one may have only a small effect.

To handle this, we first need to learn how much each marker matters. We do this using a training population, a group of individuals for which we know both the marker genotypes and the actual measured traits. A statistical model studies the training population and estimates the effect of each marker within it.

Once the model is trained, we can apply it to a population that has genotype data only. The model adds up the small effects of all its markers into a single score. This score is called the genomic estimated breeding value (GEBV), and it is what we use to rank and select candidates.

With a genome-based approach, we no longer have to wait for a trait to appear, slaughter an animal, or run a destructive disease test. We can read an individual’s DNA at birth, obtain its SNP genotypes, and predict its future performance. That prediction then becomes the basis for selection.

Why genomic selection matters

Genomic selection is now used across many food industries, including livestock, poultry, crops, and aquaculture.

The effect is real. Selective breeding has improved these industries for many decades, and genomic selection has accelerated that progress in recent years. Dairy cattle adopted it in the late 2000s, and aquaculture species such as salmon followed later. As a result, production cycles have shortened and growth performance has improved.

It is also becoming available to smaller producers. The cost of genotyping an individual has fallen over time, which puts genomic selection within reach of industries that could not afford it before.

You might feel overwhelmed by the many technical terms I mentioned earlier. However, do not worry; the fundamentals of genomics-based selective breeding can be learned step by step. Start with the biological foundation below.

1. Biological fundamentals

This section introduces the biological concepts that connect inheritance, genetic variation, markers, and traits.

2. Mathematical fundamentals

This section develops the statistical and mathematical ideas needed to understand how genomic prediction models learn from phenotypes and markers.

  1. Descriptive statistics
  2. Probability and random variables
  3. Probability distributions
  4. Sampling and statistical uncertainty
  5. Covariance and correlation
  6. Vectors and matrices
  7. Linear regression
  8. Multiple regression and regularisation
  9. Statistical models and estimation
  10. Linear mixed models

3. Genomic prediction

This section connects quantitative genetics with the marker-based models used to predict breeding values and rank selection candidates.

  1. Quantitative genetics: from phenotype to breeding value
  2. From markers to genomic prediction
  3. Genomic data encoding
  4. Genomic relationships