Rated 4.98-stars across 4.8K+ reviews
Rated 4.98-stars across 3.9K+ reviews Rated 4.98-stars across 3.9K+ reviews Rated 4.98-stars across 3.9K+ reviews Rated 4.98-stars across 3.9K+ reviews Rated 4.98-stars across 3.9K+ reviews

GWAS Explained: How Scientists Find Gene-Trait Associations

Dr. Kaet (Lukkaet Laoprapaipan) profile image By
Dr. Kaet (Lukkaet Laoprapaipan)
|
Aug 31, 2026
|
54
Genetics
Research
GWAS explained
Summary
GWAS explained

GWAS is how scientists scan the genomes of hundreds of thousands of people to find genetic spots linked to disease or traits. This article explains the method, Manhattan plots, and the key limitations to keep in mind.

Key Points in 1 Minute

  • GWAS (Genome-Wide Association Studies) scan the genomes of large numbers of people to find genetic positions that appear "more often" in people who share a disease or trait.
  • What a GWAS compares are SNPs (single spots where the genetic code can differ by one letter) between a case group and a control group.
  • A Manhattan plot is the classic way to summarise GWAS results — tall spikes mark genome positions significantly associated with the trait being studied.
  • The crucial caveat: association is not causation — a SNP that lights up is not necessarily the "culprit" behind a disease.
  • Most GWAS databases are still dominated by people of European ancestry, so results can be less accurate when applied to Asian or other ancestry groups.

Hello, everyone. As a doctor who works with genetic data every day, I'd like to introduce you to a tool that sits behind all those headlines you often see — "scientists find gene linked to disease X." That tool is called GWAS. It is a cornerstone of modern genetics and the source of the genetic risk scores we use today. At the same time, it is a tool that is very easily misunderstood, so I want to explain both its strengths and its limitations as honestly as I can.

How Does a GWAS Work?

Imagine we have two groups of people. The first group has type 2 diabetes; the second does not (we call them controls). We then read everyone's DNA at the same millions of positions across the whole genome and ask a simple question: "Are there any positions where one genetic pattern shows up in the disease group significantly more often than in the control group?" That, in a nutshell, is what a GWAS does.

The positions we compare are called SNPs (Single Nucleotide Polymorphisms) — spots where a person's genetic code can differ by just a single letter, for example some people carry an A while others carry a G. Humans have millions of common SNPs, and a GWAS systematically scans for an association one position at a time. The rough steps are:

  • Collect a large sample: A solid GWAS usually needs tens of thousands to hundreds of thousands of people — the more participants, the easier it is to see weak signals clearly.
  • Read the genotypes: Use genotyping chips or sequencing to read SNPs across everyone's genome.
  • Test statistically: Compare the frequency of each SNP between the two groups and calculate a p-value.
  • Adjust for confounders: Control for factors such as age, sex, and ancestry to reduce false positives.

Manhattan Plots: Reading the Signature GWAS Chart

If you have ever seen a chart with dense dots strung out in a long line, with tall spikes rising up here and there like the skyline of a city full of skyscrapers, that is a Manhattan plot. The name comes from its resemblance to the Manhattan skyline. Here is how to read it:

  • X-axis: The position along the genome, ordered from chromosome 1 through to the sex chromosomes.
  • Y-axis: The value of −log₁₀(p-value) — simply put, the higher a dot sits, the more statistically significant the association.
  • Threshold line: Researchers usually draw a line at a p-value of about 5×10⁻⁸, the standard for genome-wide significance. Points that rise above this line are considered credible enough to take seriously.

The reason the threshold is as strict as 5×10⁻⁸ is that we test millions of SNPs at once. If we used a loose everyday cut-off like p < 0.05, we would flag tens of thousands of "false positives" purely by chance.

Why Association Is Not Causation

This is the point I always stress with my patients. When a GWAS says a SNP is "associated" with a disease, it does not mean that SNP directly "causes" the disease, for several reasons:

  1. Linkage disequilibrium: The SNP that lights up is often not the true culprit but merely a "neighbour" inherited alongside the real causal variant sitting nearby on the chromosome.
  2. Small effect sizes: Most SNPs that GWAS find raise risk only slightly. Complex diseases like diabetes or heart disease arise from hundreds to thousands of SNPs combined, plus environment and lifestyle.
  3. Ancestry as a confounder: If population structure is not controlled well, ancestry differences can create spurious associations that have nothing to do with the biology of the disease.

So carrying one risk SNP is not a "verdict" that you will develop a disease — it only shifts your probability slightly. This idea is the very foundation of polygenic risk scores, which combine the weak signals from many SNPs into a single score, and it also underpins how we assess genetic cancer risk.

The Ancestry Diversity Gap: A Limitation We Must Talk About

The limitation I consider most important in GWAS today is the lack of diversity in the databases. Most of the world's GWAS data (estimated at more than 80% across several periods) comes from participants of European ancestry, even though this group is only a minority of the global population. This has real consequences for Asians like us, in several ways:

  • SNP frequencies and linkage disequilibrium patterns differ across ancestries, so results from one population may be inaccurate when applied to another.
  • Risk scores developed from European data often lose accuracy noticeably when applied to people of African, Asian, or Latin American ancestry.
  • Some genetic variants that matter specifically in certain ancestries may never be discovered at all, simply because no one studied that group.

This is why the global genetics community is pushing for studies in more diverse populations, and it is why having databases built from Thai and Asian people ourselves is so valuable.

What GWAS Does NOT Tell Us

To be fair to the science, I want to spell out the limitations clearly so we do not over-interpret the results:

  • It does not reveal mechanism: GWAS points to positions that are "associated" but does not explain how the gene malfunctions. That requires further laboratory experiments.
  • It does not predict destiny: For complex diseases, genetics is only one part. Lifestyle, diet, exercise, and environment all play major, modifiable roles.
  • It does not cover rare variants: GWAS is designed to catch common variants. Rare variants with strong effects (as in some inherited disorders) usually require other methods to find.
  • It is not a diagnosis: Results from GWAS or a risk score are supporting information, not a medical diagnosis. Health decisions should always involve a doctor.

In my view, GWAS is a powerful tool that has transformed medicine. But its real value only emerges when we use it with an understanding of its limits — and do not read our genes as an unchangeable fate.

1. How is a GWAS different from an ordinary DNA test?

A GWAS is population-level research that compares the SNPs of tens to hundreds of thousands of people to discover genetic positions associated with a disease or trait. An individual DNA test, by contrast, applies the knowledge gained from GWAS to estimate one person's risk. Put simply, GWAS is the discovery, and the test is the application.

2. If I carry a risk SNP that GWAS found, does that mean I will definitely get the disease?

No. Most SNPs that GWAS find raise risk only slightly. Complex diseases arise from many genetic positions combined with lifestyle and environment. Carrying a risk SNP only shifts your probability a little, not a verdict, and many factors remain modifiable through healthy habits.

3. What can a Manhattan plot tell us?

A Manhattan plot summarises a whole-genome GWAS in a single image. The X-axis is position along the chromosomes and the Y-axis is the level of statistical significance. Spikes that rise above the threshold line (around 5x10 to the minus 8) mark positions credibly associated with the trait being studied.

4. Why does ancestry diversity matter for GWAS?

Because most GWAS data comes from people of European ancestry, the risk scores built from it often become less accurate when applied to Asian or other ancestry groups. Building databases from Thai and Asian populations ourselves helps make risk assessment more accurate and fair for everyone.

References

  1. National Human Genome Research Institute (NHGRI). Genome-Wide Association Studies Fact Sheet. genome.gov
  2. Uffelmann E, et al. Genome-wide association studies. Nature Reviews Methods Primers. 2021. nature.com
  3. Sirugo G, Williams SM, Tishkoff SA. The Missing Diversity in Human Genetic Studies. Cell. 2019. cell.com
  4. MedlinePlus (NIH). What are single nucleotide polymorphisms (SNPs)?. medlineplus.gov
Written by Dr. Kaet (Lukkaet Laoprapaipan)
chat whatsapp chat line chat facebook