Skip to content
Tech News
← Back to articles

Within-family effect of ancestry on complex traits in a Mexican population

read original get Who We Are and How We Got Here (David Reich) → more articles
Why This Matters

This study uses the Mexico City Prospective Study — a cohort of over 150,000 adults — plus 1KG and HGDP reference data to measure how genetic ancestry relates to complex traits within families, a design that helps separate genetic effects from social and environmental confounding. It matters because genomics research and the polygenic tools built on it have been overwhelmingly based on European-ancestry cohorts, leaving admixed Latin American populations underserved. Large, well-characterized non-European biobanks like MCPS are the raw material for making precision medicine work for more people.

Key Takeaways
Worth a Look

Who We Are and How We Got Here (David Reich) — If this Mexico City cohort study of ancestry and complex traits sparked your curiosity, David Reich's book is the perfect companion read on how admixture and ancient migrations shaped modern populations. It explains the population-genetics concepts behind terms like admixture proportions and superpopulations in plain, gripping language.

See Who We Are and How We Got Here (David Reich) on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

Study population and ethics declaration

The MCPS is a prospective cohort of over 150,000 adult participants from Mexico City3,4. The baseline survey took place between 14 April 1998 and 28 September 2004 and focused on households in two urban districts, Coyoacán and Iztapalapa. Residents aged 35 years or older were invited to participate in the study. Of the 112,333 households with eligible residents, at least one person from 106,059 households participated3,4.

Admixture analysis and estimation of genome-wide ancestry proportions

Whole-genome sequence data from the 1KG and the HGDP were downloaded and filtered to the set of autosomal variants present on the Illumina Global Screening Array v.2 (GSAv.2). These datasets were then merged with the MCPS GSAv.2 array dataset with 140,829 participants (previously quality controlled as described in ref. 4 with the adjustment of filtering for genotype missingness before filtering for individual missingness, allowing us to retain 2,318 more participants than the 138,511 people evaluated in that study). This resulted in a merged dataset of 485,043 unambiguous bi-allelic SNPs. 1KG and HGDP participants representing four global ‘superpopulations’ were designated as reference samples and included 765, 727, 658 and 408 people of AFR, EAS, EUR and IAM ancestry, respectively. This reference set was then supplemented with 1,000 randomly selected MCPS participants unrelated to the fourth degree, resulting in a total of 1,408 IAM and MCPS samples and 3,558 reference samples overall.

Ancestry-specific allele frequencies and per-person ancestry proportions were estimated with ADMIXTURE13 v.1.3.0. An admixture model was fit for K = 4 ancestral populations inferred among the set of reference participants using an unsupervised procedure. The choice of setting K = 4 was guided by previous work that showed that the continental-level ‘superpopulations’ most represented by genetic similarity in the MCPS cohort were IAM, EUR and, to a lesser extent, AFR and EAS4. Ancestry proportions for the remaining set of 139,829 MCPS participants were then estimated by projection with the -P option. Each of the estimated K ancestries was assigned to a global ‘superpopulation’ based on averages within the reference samples (for example, the ancestry with the highest average proportion among EUR reference samples was assigned to an inferred EUR ancestral population). Ancestry proportions from each of the four inferred ancestral populations were used in subsequent analyses.

Selection of population and family samples

From the 140,829 participants we selected two non-overlapping subsets, a set of families (consisting of two or more genotyped siblings and zero, one or two genotyped parents; Extended Data Table 1) and a separate set of people who were not related up to the fourth degree of relatedness. Pair-wise relatedness for all samples was inferred with KING52, using the option --ibdseg, as described previously4. From the 30,450 sibling pairs identified by KING, those identified as full siblings were grouped into family units. To ensure that all pairs within a family were full siblings according to KING’s estimation, 25 people were removed. Furthermore, four people were excluded to ensure that each family had no more than two putative parents. The filtration resulted in a family dataset comprising 30,407 full-sibling pairs within 16,187 families (n = 38,274). To refine the family structure further, we incorporated parent–offspring relationships, identifying people with genotype data for both parents, resulting in 1,440 complete parent–offspring trios. Consequently, the family-based dataset comprised 39,714 people with either full siblings or two genotyped parents. The ‘population sample’ was selected from the 63,130 people who were unrelated up to the fourth degree of relatedness, excluding family set people and their parents, leading to a sample set of 52,583. Restriction to this unrelated set was done to minimize confounding due to pedigree relatedness in estimating population-level effects.

Estimation of IBD sharing among siblings

For each identified family, genome-wide IBD was estimated with snipar (single-nucleotide imputation of parents)53 based on genotype array data. snipar employs a hidden Markov model to infer IBD segments shared between siblings, achieving near-theoretical accuracy and reducing IBD errors compared to KING53. When running snipar, genotyping array variants were filtered based on the following criteria: minor allele frequency less than 5%, significant deviation from Hardy–Weinberg equilibrium (P < 1 × 10−4), or missingness greater than 1%. Furthermore, a genotyping error probability threshold of 4.5 × 10−4 was applied.

A comparison of estimated genome-wide IBD proportions among sibling pairs using either genetic or physical genome length is shown in Supplementary Fig. 11. For subsequent analyses, we used the realized relationships among sibling pairs estimated from snipar, with IBD defined as the genome-wide proportion of genetic map length (in centimorgans) shared.

... continue reading