{"id":460,"date":"2026-08-24T17:52:26","date_gmt":"2026-08-24T17:52:26","guid":{"rendered":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/?post_type=chapter&#038;p=460"},"modified":"2026-08-26T13:52:10","modified_gmt":"2026-08-26T13:52:10","slug":"genomic-variation-and-polygenic-scores","status":"publish","type":"chapter","link":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/chapter\/genomic-variation-and-polygenic-scores\/","title":{"raw":"Genomic Variation and Polygenic Scores","rendered":"Genomic Variation and Polygenic Scores"},"content":{"raw":"<h2>Reading Objectives<\/h2>\r\n<h2>Key Terms<\/h2>\r\n<h2>Introduction<\/h2>\r\nTERMS\r\n\r\nGWAS\r\n\r\nPGS SCORE\r\n\r\nrecall: phenotype\r\n<h2>2. From a biospecimen to genomic data<\/h2>\r\nChapter 4 investigated inherited variation without measuring specific DNA differences. This section follows how a biological specimen becomes genomic data, which is the used for genome-wide association studies (GWAS) and polygenic scores (PGS).\r\n<h3>2.1 The minimum genomic vocabulary<\/h3>\r\nDNA and phenotype were introduced in Chapter 4. Figure 5.1 introduces the additional vocabulary needed to understand how researchers represent genetic variation in GWAS and polygenic scores. Read the figure as a sequence of increasingly focused views: from the genome, to a chromosome, to a particular locus, and finally to the alleles and genotypes recorded at one SNP.\r\n\r\n[caption id=\"attachment_463\" align=\"aligncenter\" width=\"1672\"]<img class=\"wp-image-463 size-full\" src=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_33_06-PM.png\" alt=\"\" width=\"1672\" height=\"941\" \/> Figure 5.1. From genome to allele dosage. The figure follows one A\/G variant from its genomic location to three possible participant genotypes. When G is designated the effect allele, AA, AG, and GG contain 0, 1, and 2 copies of that allele.[\/caption]\r\n\r\nThe designation <strong>effect allele<\/strong> identifies which allele a reported GWAS estimate describes. It does not mean that the allele is harmful or that it causes the measured outcome. The same allele can also be associated with higher values in one analysis and lower values in another, depending on how the outcome and estimate are defined.\r\n\r\nA variant may occur inside or outside a gene. Even when a variant is located within or near a gene, its location alone does not establish that the gene explains an association. Addiction-related outcomes are related to thousands of genetic differences together with development and lived conditions.\r\n<h3>2.2 The physical and operational journey<\/h3>\r\nGenomic data collection begins with a youth and caregiver complete assent and consent procedures for sample collection and genomic sequencing (assent is the youth\u2019s voluntary, age\u2011appropriate agreement to participate, while consent is the caregiver\u2019s legally valid authorization). Staff then collect saliva or blood using a documented protocol. The original ABCD protocol described site-level ethics approval, participant consent and assent, baseline saliva collection, repository shipment, storage, DNA isolation, and blood collection for part of the twin sample (Uban et al., 2018).\r\n\r\nThe specimen receives a controlled sample identifier linked to the participant record. The two identifiers serve different purposes: a participant identifier refers to a person, while a sample identifier refers to a particular tube, collection, or extracted portion.\u00a0The specimen is stored, transported, and processed under documented conditions. Laboratory staff isolate DNA, assess its quantity and quality, and use a genotyping array or sequencing instrument to generate laboratory files. Quality checks may identify contamination, low-quality results, unexpected relatedness, or disagreement between records and genetic data. Files then move through secure systems for additional processing, documentation, and linkage to approved outcome measurements.\r\n\r\nThis sequence explains why laboratory output is not immediately an analysis-ready dataset. Each stage can change what is available for analysis. Good provenance records the specimen source, data release, laboratory platform, processing pipeline, reference genome, quality-control decisions, and software or file version.\r\n\r\n[caption id=\"attachment_464\" align=\"aligncenter\" width=\"1672\"]<img class=\"wp-image-464 size-full\" src=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_35_39-PM.png\" alt=\"\" width=\"1672\" height=\"941\" \/> Figure 5.2. The genomic biomedical data journey. Five stages, collect, generate, curate, analyze, and report or reuse, are shown in three aligned bands. The first band follows people, specimens, laboratories, and computing systems. The second follows participant and sample identifiers, laboratory files, genomic records, GWAS statistics, and participant-level PGS variables. The third follows consent and assent, secure linkage, provenance, controlled access, approved analysis, and responsible reporting.[\/caption]\r\n<h3>2.3 Genomic resources in ABCD<\/h3>\r\nABCD provides several related genomic resources. <strong>Genotyping-array data<\/strong> record selected variants measured using probes on an array. <strong>Whole-genome sequencing data<\/strong> are produced by reading DNA across the genome and aligning the reads to a reference sequence. Both resources require quality review.\r\n\r\n<strong>Imputed genotype data<\/strong> add statistically estimated genotypes at variants that were not directly measured on the array. The estimates use observed genotypes and patterns in a reference resource. An imputed allele dosage can therefore be a value between 0 and 2, reflecting uncertainty, rather than a directly observed count. Imputation expands genomic coverage, but it does not mean that every DNA position was directly observed.\r\n\r\nhttps:\/\/www.youtube.com\/watch?v=MvuYATh7Y74\r\n\r\nABCD also provides sample-level and variant-level quality information, estimates of genetic relatedness, and <strong>principal components<\/strong>, which are numeric variables summarizing major patterns of genomic similarity among participants. Relatedness measures help identify family relationships and account for nonindependence. Principal components help researchers address population structure that could otherwise create misleading associations.\r\n<blockquote><strong>ABCD Release 7.0 note.<\/strong> Release 7.0 includes curated Smokescreen array data for 11,670 participants at about 515,000 variants, imputed array data for the same participants at about 260 million variants, and 30x Illumina whole-genome sequencing for 8,710 participants at about 169 million variants. Counts depend on the resource and quality-control rules and should not be treated as permanent study characteristics. Record the release, retrieval date, platform, and analytic sample whenever reporting results. (<a href=\"https:\/\/docs.abcdstudy.org\/latest\/documentation\/non_imaging\/gn.html\">ABCD genetics documentation<\/a>; <a href=\"https:\/\/docs.abcdstudy.org\/latest\/documentation\/release_notes\/7_0.html\">Release 7.0 notes<\/a>)<\/blockquote>\r\n<h3>2.4 One resource, several data structures<\/h3>\r\nA conceptual genotype matrix has participants in rows and variants in columns, with each cell containing a genotype or dosage. Real genomic resources are too large and specialized to remain in one ordinary DataFrame. VCF files can store variant and genotype information, while PLINK files provide another specialized representation used in genetic analysis. We do not work with these files in this course - but will work synthetic versions of derived data.\r\n\r\nLater steps create different structures for different purposes. GWAS summary statistics usually contain one row per tested variant, along with its effect allele, estimated association, and uncertainty. A PGS scoring file contains the variants, effect alleles, and weights used to construct a score. An analysis-ready participant table may then contain one standardized PGS per participant, together with selected outcomes and contextual variables. The PGS is therefore a derived variable, not raw biological data.\r\n<h3>2.5 Governance throughout the journey<\/h3>\r\nRemoving names does not eliminate genomic privacy risks. A person\u2019s genomic data can be highly identifying, and they are relational because they can also reveal information about biological relatives. Data collected from youth may later support research questions that participants and families did not anticipate at enrollment. Linking specimens, participant records, genomic files, and outcomes therefore requires controlled identifiers, approved systems, and continuing oversight.\r\n\r\nControlled access is more than a password. Review may consider the investigator and institution, the proposed purpose, the computing environment, the specific data requested, plans for storage and sharing, reporting procedures, deletion requirements, and compliance with prior agreements. These safeguards make authority, purpose, and accountability explicit.\r\n<h2>3. From a measured outcome to GWAS evidence<\/h2>\r\nSection 2 described how researchers sequence a genome, in other words, how genetic data is made and what are its data structures. In Section 3 we look at genetic variations are associated with phenotypes, in other words, how do we know if our genes make us predisposed for certain physical or behavioral traits?\r\n<p class=\"PDq2pG_selectionAnchorContainer\" data-start=\"601\" data-end=\"906\">A <a href=\"https:\/\/www.broadinstitute.org\/visuals\/explainer-genome-wide-association-studies\"><strong data-start=\"603\" data-end=\"643\">genome-wide association study (GWAS)<\/strong><\/a> examines genetic variants across the genome to determine whether any are statistically associated with a measured outcome. Researchers study many participants and test hundreds of thousands or millions of variants, usually single-nucleotide polymorphisms (SNPs).<\/p>\r\n<p data-start=\"911\" data-end=\"1194\">The illustration below shows the basic logic using a binary disease outcome. Researchers compare allele frequencies between participants with and without the disease. If an allele is more frequent in one group, that variant may show evidence of association with the measured outcome.<\/p>\r\n\r\n\r\n[caption id=\"attachment_132\" align=\"aligncenter\" width=\"1024\"]<img class=\"wp-image-132 size-large\" src=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/02\/Gwas_NHGRI-1024x576.jpg\" alt=\"Diagram illustrating a genome-wide association study (GWAS), comparing people with a disease vs without a disease across multiple SNPs, showing one SNP associated with disease.\" width=\"1024\" height=\"576\" \/> <strong data-start=\"1241\" data-end=\"1288\">Figure 5.3. A simplified case\u2013control GWAS.<\/strong> Researchers compare allele frequencies between participants with and without a disease. In this example, SNPs 1 and 2 show no association, while SNP 3 differs between the groups. This is one type of GWAS; studies can also test continuous measures and counts. An association does not by itself establish that a variant causes the outcome. Adapted from <a class=\"decorated-link\" href=\"https:\/\/www.genome.gov\/about-genomics\/fact-sheets\/Genome-Wide-Association-Studies-Fact-Sheet\" target=\"_new\" rel=\"noopener\" data-start=\"1640\" data-end=\"1741\">NHGRI<\/a>.[\/caption]\r\n<p data-start=\"1747\" data-end=\"2102\">This picture provides the basic idea, but it leaves out an essential question: What exactly counts as \u201cwith\u201d or \u201cwithout\u201d the outcome? Addiction research may instead examine frequency of use, number of use days, symptom counts, or a broader behavioral score. The measurement decisions made before the genetic analysis determine what the GWAS results mean.<\/p>\r\n\r\n<h3>3.1 The measured outcome comes first<\/h3>\r\nChapters 2 through 4 followed the measurement process:\r\n\r\n<strong>Construct \u2192 instrument or procedure \u2192 recorded value \u2192 scoring process \u2192 analytic variable \u2192 interpretation<\/strong>\r\n\r\nThe same sequence applies in a GWAS. Researchers must specify the phenotype, the instrument or procedure used to measure it, and the analytic variable that will represent it. \u201cA GWAS of addiction\u201d is not sufficiently precise.\r\n\r\nFor example, researchers could examine whether someone has ever used cannabis, past-year cannabis-use frequency, cannabis use disorder, alcohol consumption, alcohol-related problems, or a broad externalizing score. These outcomes may differ in respondent, reference period, scoring rules, developmental period, population, and degree of impairment. They answer different research questions and may produce different genomic associations.\r\n<h3>3.2 The discovery sample and GWAS process<\/h3>\r\nThe <strong>discovery sample<\/strong> consists of the participants whose genomic and outcome data are used to estimate associations. Researchers define eligibility rules, identify qualifying records, apply quality controls, account for relevant study characteristics and population structure, and test variants across the genome.\r\n\r\nIn a case-control GWAS, researchers ask whether genotypes differ systematically between participants who meet specified criteria and those who do not. GWAS can also examine continuous measures, such as an externalizing score, or counts, such as reported days of cannabis use. The statistical model must match the outcome.\r\n\r\nGWAS <strong>summary statistics<\/strong> commonly record each variant, its genomic position and effect allele, its estimated association, uncertainty, and statistical evidence. They summarize results across the discovery sample rather than providing one row per participant.\r\n<h3>3.3 Quality control and population structure<\/h3>\r\nQuality checks may identify incomplete records, poorly measured variants, duplicate samples, unexpected biological relationships, differences across laboratory batches or platforms, and low-quality imputed genotypes. These checks can change both the analytic sample and the variants tested.\r\n\r\nResearchers must also consider <strong>population structure<\/strong>, meaning patterned genomic similarity related to population histories. Without appropriate adjustment, an analysis may detect background population differences rather than the relationship of interest. <strong>Genetic principal components<\/strong> summarize some major patterns of genomic similarity and are often included in GWAS models. They do not represent natural divisions of people into genetic groups.\r\n<h3>3.4 Testing many variants and seeking replication<\/h3>\r\nA GWAS evaluates hundreds of thousands or millions of variant-outcome associations. Conducting so many tests creates many opportunities for apparently strong results to occur by chance. Researchers therefore use a stringent <strong>genome-wide significance threshold<\/strong> as one safeguard.\r\n\r\nCrossing the threshold indicates strong statistical evidence under the model. It does not mean that the association is large, clinically useful, causal, or biologically understood. Calculating p-values and multiple-testing corrections is reserved for DSARM 2. Here, the goal is interpreting the threshold.\r\n\r\nA <strong>replication sample<\/strong> is a separate sample used to examine whether an important association appears again. Replication reduces the possibility that a finding reflects chance or a feature unique to one sample. It strengthens confidence in an association but does not prove a biological mechanism.\r\n<h3>3.5 Reading a Manhattan plot<\/h3>\r\nA <strong>Manhattan plot<\/strong> displays GWAS results across the genome. The horizontal axis shows genomic position, grouped by chromosome. Each point represents one tested variant. The vertical axis represents the strength of statistical evidence for its association with the measured outcome, so higher points indicate stronger evidence. A horizontal line marks the genome-wide significance threshold.\r\n\r\nSeveral high points may appear together because of <strong>linkage disequilibrium<\/strong>, meaning that particular alleles at nearby locations tend to be inherited together. One measured variant may therefore tag another in the same region. A peak usually identifies an associated genomic region, not one proven causal nucleotide. Assigning the peak to the nearest gene is an annotation step, not proof that the gene explains the association.\r\n\r\nA careful interpretation is: <em>This region shows statistical evidence of association with the measured outcome in the discovery sample.<\/em> The plot alone does not identify a causal variant, gene, biological pathway, or intervention.\r\n\r\n[caption id=\"attachment_470\" align=\"aligncenter\" width=\"1672\"]<img class=\"wp-image-470 size-full\" src=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-26-2026-09_47_14-AM.png\" alt=\"\" width=\"1672\" height=\"941\" \/> <strong>Figure 5.4<\/strong>. From outcome definition to a Manhattan plot. Panel A traces the path from outcome definition through the discovery sample, quality control, variant tests, and summary statistics. Panel B shows a synthetic Manhattan plot for past-year cannabis-use frequency. A high point indicates statistical evidence of association at a genomic location. It does not identify a causal variant, gene, mechanism, or intervention.[\/caption]\r\n<h3>3.6 Running example and research extension<\/h3>\r\nThe synthetic example in Figure 5.3 uses past-year cannabis-use frequency as the measured outcome. Every point represents a test of one variant against that same defined variable in the qualifying discovery sample. Because the data are synthetic, the peak illustrates how to read the plot and does not make a claim about an actual genomic region.\r\n<blockquote><strong>Research extension: Multivariate GWAS<\/strong>\r\n\r\nSome studies model several related outcomes together. A 2026 study combined substance-use-disorder and other externalizing GWAS representing more than 2.2 million participants. The researchers examined associations shared across broad externalizing behavior and associations more specific to particular substance-use disorders. Combining outcomes can improve detection of small associations, but it changes what the analysis represents. Large sample size also does not guarantee broad population coverage. The study included only participants classified as having European genetic ancestry, limiting the populations to which its findings can be generalized (<a href=\"https:\/\/www.nature.com\/articles\/s44220-026-00608-6\">Poore et al., 2026<\/a>).<\/blockquote>\r\nThe summary statistics produced by a GWAS are the endpoint of this section and the starting point for the next one. Section 4 explains how researchers combine selected variant estimates with participant genotypes to construct a polygenic score.","rendered":"<h2>Reading Objectives<\/h2>\n<h2>Key Terms<\/h2>\n<h2>Introduction<\/h2>\n<p>TERMS<\/p>\n<p>GWAS<\/p>\n<p>PGS SCORE<\/p>\n<p>recall: phenotype<\/p>\n<h2>2. From a biospecimen to genomic data<\/h2>\n<p>Chapter 4 investigated inherited variation without measuring specific DNA differences. This section follows how a biological specimen becomes genomic data, which is the used for genome-wide association studies (GWAS) and polygenic scores (PGS).<\/p>\n<h3>2.1 The minimum genomic vocabulary<\/h3>\n<p>DNA and phenotype were introduced in Chapter 4. Figure 5.1 introduces the additional vocabulary needed to understand how researchers represent genetic variation in GWAS and polygenic scores. Read the figure as a sequence of increasingly focused views: from the genome, to a chromosome, to a particular locus, and finally to the alleles and genotypes recorded at one SNP.<\/p>\n<figure id=\"attachment_463\" aria-describedby=\"caption-attachment-463\" style=\"width: 1672px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-463 size-full\" src=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_33_06-PM.png\" alt=\"\" width=\"1672\" height=\"941\" srcset=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_33_06-PM.png 1672w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_33_06-PM-300x169.png 300w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_33_06-PM-1024x576.png 1024w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_33_06-PM-768x432.png 768w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_33_06-PM-1536x864.png 1536w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_33_06-PM-65x37.png 65w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_33_06-PM-225x127.png 225w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_33_06-PM-350x197.png 350w\" sizes=\"auto, (max-width: 1672px) 100vw, 1672px\" \/><figcaption id=\"caption-attachment-463\" class=\"wp-caption-text\">Figure 5.1. From genome to allele dosage. The figure follows one A\/G variant from its genomic location to three possible participant genotypes. When G is designated the effect allele, AA, AG, and GG contain 0, 1, and 2 copies of that allele.<\/figcaption><\/figure>\n<p>The designation <strong>effect allele<\/strong> identifies which allele a reported GWAS estimate describes. It does not mean that the allele is harmful or that it causes the measured outcome. The same allele can also be associated with higher values in one analysis and lower values in another, depending on how the outcome and estimate are defined.<\/p>\n<p>A variant may occur inside or outside a gene. Even when a variant is located within or near a gene, its location alone does not establish that the gene explains an association. Addiction-related outcomes are related to thousands of genetic differences together with development and lived conditions.<\/p>\n<h3>2.2 The physical and operational journey<\/h3>\n<p>Genomic data collection begins with a youth and caregiver complete assent and consent procedures for sample collection and genomic sequencing (assent is the youth\u2019s voluntary, age\u2011appropriate agreement to participate, while consent is the caregiver\u2019s legally valid authorization). Staff then collect saliva or blood using a documented protocol. The original ABCD protocol described site-level ethics approval, participant consent and assent, baseline saliva collection, repository shipment, storage, DNA isolation, and blood collection for part of the twin sample (Uban et al., 2018).<\/p>\n<p>The specimen receives a controlled sample identifier linked to the participant record. The two identifiers serve different purposes: a participant identifier refers to a person, while a sample identifier refers to a particular tube, collection, or extracted portion.\u00a0The specimen is stored, transported, and processed under documented conditions. Laboratory staff isolate DNA, assess its quantity and quality, and use a genotyping array or sequencing instrument to generate laboratory files. Quality checks may identify contamination, low-quality results, unexpected relatedness, or disagreement between records and genetic data. Files then move through secure systems for additional processing, documentation, and linkage to approved outcome measurements.<\/p>\n<p>This sequence explains why laboratory output is not immediately an analysis-ready dataset. Each stage can change what is available for analysis. Good provenance records the specimen source, data release, laboratory platform, processing pipeline, reference genome, quality-control decisions, and software or file version.<\/p>\n<figure id=\"attachment_464\" aria-describedby=\"caption-attachment-464\" style=\"width: 1672px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-464 size-full\" src=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_35_39-PM.png\" alt=\"\" width=\"1672\" height=\"941\" srcset=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_35_39-PM.png 1672w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_35_39-PM-300x169.png 300w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_35_39-PM-1024x576.png 1024w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_35_39-PM-768x432.png 768w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_35_39-PM-1536x864.png 1536w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_35_39-PM-65x37.png 65w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_35_39-PM-225x127.png 225w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-24-2026-02_35_39-PM-350x197.png 350w\" sizes=\"auto, (max-width: 1672px) 100vw, 1672px\" \/><figcaption id=\"caption-attachment-464\" class=\"wp-caption-text\">Figure 5.2. The genomic biomedical data journey. Five stages, collect, generate, curate, analyze, and report or reuse, are shown in three aligned bands. The first band follows people, specimens, laboratories, and computing systems. The second follows participant and sample identifiers, laboratory files, genomic records, GWAS statistics, and participant-level PGS variables. The third follows consent and assent, secure linkage, provenance, controlled access, approved analysis, and responsible reporting.<\/figcaption><\/figure>\n<h3>2.3 Genomic resources in ABCD<\/h3>\n<p>ABCD provides several related genomic resources. <strong>Genotyping-array data<\/strong> record selected variants measured using probes on an array. <strong>Whole-genome sequencing data<\/strong> are produced by reading DNA across the genome and aligning the reads to a reference sequence. Both resources require quality review.<\/p>\n<p><strong>Imputed genotype data<\/strong> add statistically estimated genotypes at variants that were not directly measured on the array. The estimates use observed genotypes and patterns in a reference resource. An imputed allele dosage can therefore be a value between 0 and 2, reflecting uncertainty, rather than a directly observed count. Imputation expands genomic coverage, but it does not mean that every DNA position was directly observed.<\/p>\n<p><iframe loading=\"lazy\" id=\"oembed-1\" title=\"How to sequence the human genome - Mark J. Kiel\" width=\"500\" height=\"281\" src=\"https:\/\/www.youtube.com\/embed\/MvuYATh7Y74?feature=oembed&#38;rel=0&#38;enablejsapi=1&#38;origin=https:\/\/openpub.libraries.rutgers.edu\" frameborder=\"0\" allowfullscreen=\"allowfullscreen\"><\/iframe><\/p>\n<p>ABCD also provides sample-level and variant-level quality information, estimates of genetic relatedness, and <strong>principal components<\/strong>, which are numeric variables summarizing major patterns of genomic similarity among participants. Relatedness measures help identify family relationships and account for nonindependence. Principal components help researchers address population structure that could otherwise create misleading associations.<\/p>\n<blockquote><p><strong>ABCD Release 7.0 note.<\/strong> Release 7.0 includes curated Smokescreen array data for 11,670 participants at about 515,000 variants, imputed array data for the same participants at about 260 million variants, and 30x Illumina whole-genome sequencing for 8,710 participants at about 169 million variants. Counts depend on the resource and quality-control rules and should not be treated as permanent study characteristics. Record the release, retrieval date, platform, and analytic sample whenever reporting results. (<a href=\"https:\/\/docs.abcdstudy.org\/latest\/documentation\/non_imaging\/gn.html\">ABCD genetics documentation<\/a>; <a href=\"https:\/\/docs.abcdstudy.org\/latest\/documentation\/release_notes\/7_0.html\">Release 7.0 notes<\/a>)<\/p><\/blockquote>\n<h3>2.4 One resource, several data structures<\/h3>\n<p>A conceptual genotype matrix has participants in rows and variants in columns, with each cell containing a genotype or dosage. Real genomic resources are too large and specialized to remain in one ordinary DataFrame. VCF files can store variant and genotype information, while PLINK files provide another specialized representation used in genetic analysis. We do not work with these files in this course &#8211; but will work synthetic versions of derived data.<\/p>\n<p>Later steps create different structures for different purposes. GWAS summary statistics usually contain one row per tested variant, along with its effect allele, estimated association, and uncertainty. A PGS scoring file contains the variants, effect alleles, and weights used to construct a score. An analysis-ready participant table may then contain one standardized PGS per participant, together with selected outcomes and contextual variables. The PGS is therefore a derived variable, not raw biological data.<\/p>\n<h3>2.5 Governance throughout the journey<\/h3>\n<p>Removing names does not eliminate genomic privacy risks. A person\u2019s genomic data can be highly identifying, and they are relational because they can also reveal information about biological relatives. Data collected from youth may later support research questions that participants and families did not anticipate at enrollment. Linking specimens, participant records, genomic files, and outcomes therefore requires controlled identifiers, approved systems, and continuing oversight.<\/p>\n<p>Controlled access is more than a password. Review may consider the investigator and institution, the proposed purpose, the computing environment, the specific data requested, plans for storage and sharing, reporting procedures, deletion requirements, and compliance with prior agreements. These safeguards make authority, purpose, and accountability explicit.<\/p>\n<h2>3. From a measured outcome to GWAS evidence<\/h2>\n<p>Section 2 described how researchers sequence a genome, in other words, how genetic data is made and what are its data structures. In Section 3 we look at genetic variations are associated with phenotypes, in other words, how do we know if our genes make us predisposed for certain physical or behavioral traits?<\/p>\n<p class=\"PDq2pG_selectionAnchorContainer\" data-start=\"601\" data-end=\"906\">A <a href=\"https:\/\/www.broadinstitute.org\/visuals\/explainer-genome-wide-association-studies\"><strong data-start=\"603\" data-end=\"643\">genome-wide association study (GWAS)<\/strong><\/a> examines genetic variants across the genome to determine whether any are statistically associated with a measured outcome. Researchers study many participants and test hundreds of thousands or millions of variants, usually single-nucleotide polymorphisms (SNPs).<\/p>\n<p data-start=\"911\" data-end=\"1194\">The illustration below shows the basic logic using a binary disease outcome. Researchers compare allele frequencies between participants with and without the disease. If an allele is more frequent in one group, that variant may show evidence of association with the measured outcome.<\/p>\n<figure id=\"attachment_132\" aria-describedby=\"caption-attachment-132\" style=\"width: 1024px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-132 size-large\" src=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/02\/Gwas_NHGRI-1024x576.jpg\" alt=\"Diagram illustrating a genome-wide association study (GWAS), comparing people with a disease vs without a disease across multiple SNPs, showing one SNP associated with disease.\" width=\"1024\" height=\"576\" srcset=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/02\/Gwas_NHGRI-1024x576.jpg 1024w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/02\/Gwas_NHGRI-300x169.jpg 300w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/02\/Gwas_NHGRI-768x432.jpg 768w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/02\/Gwas_NHGRI-1536x864.jpg 1536w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/02\/Gwas_NHGRI-65x37.jpg 65w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/02\/Gwas_NHGRI-225x127.jpg 225w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/02\/Gwas_NHGRI-350x197.jpg 350w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/02\/Gwas_NHGRI.jpg 1920w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><figcaption id=\"caption-attachment-132\" class=\"wp-caption-text\"><strong data-start=\"1241\" data-end=\"1288\">Figure 5.3. A simplified case\u2013control GWAS.<\/strong> Researchers compare allele frequencies between participants with and without a disease. In this example, SNPs 1 and 2 show no association, while SNP 3 differs between the groups. This is one type of GWAS; studies can also test continuous measures and counts. An association does not by itself establish that a variant causes the outcome. Adapted from <a class=\"decorated-link\" href=\"https:\/\/www.genome.gov\/about-genomics\/fact-sheets\/Genome-Wide-Association-Studies-Fact-Sheet\" target=\"_new\" rel=\"noopener\" data-start=\"1640\" data-end=\"1741\">NHGRI<\/a>.<\/figcaption><\/figure>\n<p data-start=\"1747\" data-end=\"2102\">This picture provides the basic idea, but it leaves out an essential question: What exactly counts as \u201cwith\u201d or \u201cwithout\u201d the outcome? Addiction research may instead examine frequency of use, number of use days, symptom counts, or a broader behavioral score. The measurement decisions made before the genetic analysis determine what the GWAS results mean.<\/p>\n<h3>3.1 The measured outcome comes first<\/h3>\n<p>Chapters 2 through 4 followed the measurement process:<\/p>\n<p><strong>Construct \u2192 instrument or procedure \u2192 recorded value \u2192 scoring process \u2192 analytic variable \u2192 interpretation<\/strong><\/p>\n<p>The same sequence applies in a GWAS. Researchers must specify the phenotype, the instrument or procedure used to measure it, and the analytic variable that will represent it. \u201cA GWAS of addiction\u201d is not sufficiently precise.<\/p>\n<p>For example, researchers could examine whether someone has ever used cannabis, past-year cannabis-use frequency, cannabis use disorder, alcohol consumption, alcohol-related problems, or a broad externalizing score. These outcomes may differ in respondent, reference period, scoring rules, developmental period, population, and degree of impairment. They answer different research questions and may produce different genomic associations.<\/p>\n<h3>3.2 The discovery sample and GWAS process<\/h3>\n<p>The <strong>discovery sample<\/strong> consists of the participants whose genomic and outcome data are used to estimate associations. Researchers define eligibility rules, identify qualifying records, apply quality controls, account for relevant study characteristics and population structure, and test variants across the genome.<\/p>\n<p>In a case-control GWAS, researchers ask whether genotypes differ systematically between participants who meet specified criteria and those who do not. GWAS can also examine continuous measures, such as an externalizing score, or counts, such as reported days of cannabis use. The statistical model must match the outcome.<\/p>\n<p>GWAS <strong>summary statistics<\/strong> commonly record each variant, its genomic position and effect allele, its estimated association, uncertainty, and statistical evidence. They summarize results across the discovery sample rather than providing one row per participant.<\/p>\n<h3>3.3 Quality control and population structure<\/h3>\n<p>Quality checks may identify incomplete records, poorly measured variants, duplicate samples, unexpected biological relationships, differences across laboratory batches or platforms, and low-quality imputed genotypes. These checks can change both the analytic sample and the variants tested.<\/p>\n<p>Researchers must also consider <strong>population structure<\/strong>, meaning patterned genomic similarity related to population histories. Without appropriate adjustment, an analysis may detect background population differences rather than the relationship of interest. <strong>Genetic principal components<\/strong> summarize some major patterns of genomic similarity and are often included in GWAS models. They do not represent natural divisions of people into genetic groups.<\/p>\n<h3>3.4 Testing many variants and seeking replication<\/h3>\n<p>A GWAS evaluates hundreds of thousands or millions of variant-outcome associations. Conducting so many tests creates many opportunities for apparently strong results to occur by chance. Researchers therefore use a stringent <strong>genome-wide significance threshold<\/strong> as one safeguard.<\/p>\n<p>Crossing the threshold indicates strong statistical evidence under the model. It does not mean that the association is large, clinically useful, causal, or biologically understood. Calculating p-values and multiple-testing corrections is reserved for DSARM 2. Here, the goal is interpreting the threshold.<\/p>\n<p>A <strong>replication sample<\/strong> is a separate sample used to examine whether an important association appears again. Replication reduces the possibility that a finding reflects chance or a feature unique to one sample. It strengthens confidence in an association but does not prove a biological mechanism.<\/p>\n<h3>3.5 Reading a Manhattan plot<\/h3>\n<p>A <strong>Manhattan plot<\/strong> displays GWAS results across the genome. The horizontal axis shows genomic position, grouped by chromosome. Each point represents one tested variant. The vertical axis represents the strength of statistical evidence for its association with the measured outcome, so higher points indicate stronger evidence. A horizontal line marks the genome-wide significance threshold.<\/p>\n<p>Several high points may appear together because of <strong>linkage disequilibrium<\/strong>, meaning that particular alleles at nearby locations tend to be inherited together. One measured variant may therefore tag another in the same region. A peak usually identifies an associated genomic region, not one proven causal nucleotide. Assigning the peak to the nearest gene is an annotation step, not proof that the gene explains the association.<\/p>\n<p>A careful interpretation is: <em>This region shows statistical evidence of association with the measured outcome in the discovery sample.<\/em> The plot alone does not identify a causal variant, gene, biological pathway, or intervention.<\/p>\n<figure id=\"attachment_470\" aria-describedby=\"caption-attachment-470\" style=\"width: 1672px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-470 size-full\" src=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-26-2026-09_47_14-AM.png\" alt=\"\" width=\"1672\" height=\"941\" srcset=\"https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-26-2026-09_47_14-AM.png 1672w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-26-2026-09_47_14-AM-300x169.png 300w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-26-2026-09_47_14-AM-1024x576.png 1024w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-26-2026-09_47_14-AM-768x432.png 768w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-26-2026-09_47_14-AM-1536x864.png 1536w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-26-2026-09_47_14-AM-65x37.png 65w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-26-2026-09_47_14-AM-225x127.png 225w, https:\/\/openpub.libraries.rutgers.edu:443\/wp-content\/uploads\/sites\/28\/2026\/08\/ChatGPT-Image-Aug-26-2026-09_47_14-AM-350x197.png 350w\" sizes=\"auto, (max-width: 1672px) 100vw, 1672px\" \/><figcaption id=\"caption-attachment-470\" class=\"wp-caption-text\"><strong>Figure 5.4<\/strong>. From outcome definition to a Manhattan plot. Panel A traces the path from outcome definition through the discovery sample, quality control, variant tests, and summary statistics. Panel B shows a synthetic Manhattan plot for past-year cannabis-use frequency. A high point indicates statistical evidence of association at a genomic location. It does not identify a causal variant, gene, mechanism, or intervention.<\/figcaption><\/figure>\n<h3>3.6 Running example and research extension<\/h3>\n<p>The synthetic example in Figure 5.3 uses past-year cannabis-use frequency as the measured outcome. Every point represents a test of one variant against that same defined variable in the qualifying discovery sample. Because the data are synthetic, the peak illustrates how to read the plot and does not make a claim about an actual genomic region.<\/p>\n<blockquote><p><strong>Research extension: Multivariate GWAS<\/strong><\/p>\n<p>Some studies model several related outcomes together. A 2026 study combined substance-use-disorder and other externalizing GWAS representing more than 2.2 million participants. The researchers examined associations shared across broad externalizing behavior and associations more specific to particular substance-use disorders. Combining outcomes can improve detection of small associations, but it changes what the analysis represents. Large sample size also does not guarantee broad population coverage. The study included only participants classified as having European genetic ancestry, limiting the populations to which its findings can be generalized (<a href=\"https:\/\/www.nature.com\/articles\/s44220-026-00608-6\">Poore et al., 2026<\/a>).<\/p><\/blockquote>\n<p>The summary statistics produced by a GWAS are the endpoint of this section and the starting point for the next one. Section 4 explains how researchers combine selected variant estimates with participant genotypes to construct a polygenic score.<\/p>\n","protected":false},"author":30,"menu_order":4,"template":"","meta":{"pb_show_title":"on","pb_short_title":"","pb_subtitle":"","pb_authors":[],"pb_section_license":""},"chapter-type":[],"contributor":[],"license":[],"class_list":["post-460","chapter","type-chapter","status-publish","hentry"],"part":27,"_links":{"self":[{"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/pressbooks\/v2\/chapters\/460","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/pressbooks\/v2\/chapters"}],"about":[{"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/wp\/v2\/types\/chapter"}],"author":[{"embeddable":true,"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/wp\/v2\/users\/30"}],"version-history":[{"count":9,"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/pressbooks\/v2\/chapters\/460\/revisions"}],"predecessor-version":[{"id":472,"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/pressbooks\/v2\/chapters\/460\/revisions\/472"}],"part":[{"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/pressbooks\/v2\/parts\/27"}],"metadata":[{"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/pressbooks\/v2\/chapters\/460\/metadata\/"}],"wp:attachment":[{"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/wp\/v2\/media?parent=460"}],"wp:term":[{"taxonomy":"chapter-type","embeddable":true,"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/pressbooks\/v2\/chapter-type?post=460"},{"taxonomy":"contributor","embeddable":true,"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/wp\/v2\/contributor?post=460"},{"taxonomy":"license","embeddable":true,"href":"https:\/\/openpub.libraries.rutgers.edu\/dsarm12\/wp-json\/wp\/v2\/license?post=460"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}