Cohorts and measures
We incorporated personality trait data from 46 cohorts participating in the ReGPC. Cohorts were included that provided data on genotyped participants with EUR-like and AFR-like genomes who had completed a validated multi-item personality inventory compatible with the Big Five personality trait taxonomy. Descriptive statistics for each cohort are provided in Supplementary Tables 1 and 2. Descriptions of each cohort are provided in Supplementary Note 7. A PRISMA flowchart for this meta-analysis is available at OSF (https://osf.io/sh8gr). These analyses were not pre-registered. As this meta-analysis used de-identified summary-level data, it was deemed exempt from ethics board approval by the University of Texas at Austin Institutional Review Board (STUDY00001941); ethics approval for individual cohorts that collected data and contributed to the meta-analysis is presented in Supplementary Note 7. All of the participants consented to the research, and this research was performed in accordance with all relevant guidelines and regulations.
Population-level meta-analysis
For each trait, in each cohort, we conducted population-level genome-wide association analyses among participants with EUR-like and AFR-like genomes56 across the 22 chromosomes and the X chromosome, following a standard operating procedure (https://osf.io/rsc49). In each cohort, we excluded SNPs with in-sample minor allele frequency (MAF) < 1%, call rate < 95%, deviation from Hardy–Weinberg Equilibrium (P < 1 × 10−5), poor imputation quality (INFO < 0.40) and those with poor clustering on visual inspection of intensity plots. Participants were excluded if they had low overall call rates (<95%), excess autosomal heterozygosity or homozygosity, were duplicated samples, had gender inconsistent with sex (which is often indicative of mislabelled demographic data57) or had chromosomal abnormalities. Association analyses were conducted in each cohort by regressing personality trait score on each biallelic 1000 Genomes 3v5 EUR or AFR SNP58 with PLINK2 (ref. 59), PLINK60, GCTA61, BOLT-LLM62, FastGWA63 or proprietary software, using mixed models to account for relatedness when appropriate. These analyses included controls for age, age2, birth cohort, sex and genetic principal components, to address potential lifespan and birth cohort differences in personality traits1,26. Some groups contributed separate results for potentially overlapping cohorts of participants. In these cases, we meta-analysed estimates while accounting for dependency of their estimation errors. We harmonized and directionally aligned effects using EasyQC64 and checked that effects were directionally consistent across cohorts by estimating genetic correlations with LDSC34 (conducted on HapMap3 SNPs with an INFO threshold of ≥0.90) in GenomicSEM37.
For each Big Five trait, we then conducted IVW meta-analysis of GWAS estimates across all of the contributing cohorts, separately for participants with EUR-like and AFR-like genomes, in R (v.4.3.1)65. For each meta-analysed SNP, we re-estimated n using observed meta-analytic MAF and standard error. This produced 10 sets of summary statistics (five traits × two ancestral groups). We quantified the number of genome-wide significant lead SNPs at P < 5 × 10−8 that were linkage disequilibrium (LD)-independent of other lead SNPs nearby on the genome (within a 250 kb window) using FUMA66. We also identified lead SNPs that were LD-independent of all other lead SNPs on the chromosome regardless of genomic distance. Ancestry-stratified Manhattan plots of meta-analytic genome-wide associations are presented in Extended Data Fig. 1.
We next conducted transancestry IVW meta-analyses to synthesize data across the k = 10 cohorts that included participants with both EUR-like and AFR-like genomes, and the additional k = 36 cohorts including solely participants with AFR-like genomes. As around 95% of participants were of EUR-like ancestry, and because stronger linkage disequilibrium among EUR-like compared with AFR-like superpopulations67 means that using the EUR reference panel probably leads to conservative estimation of the number of lead SNPs, we initially identified candidate lead SNPs using the 1000 Genomes 3v5 EUR reference panel. Using information on SNP LD from the 1000 Genomes 3v5 EUR and AFR reference panels, we estimated transancestry sample size-weighted LD matrices for each chromosome across all candidate transancestry lead SNPs and all lead SNPs from the ancestry-stratified analyses, which we used to identify additional LD-independent lead SNPs. To supplement transancestry IVW results, we also conducted transancestry meta-analyses using MAMA31.
We tested the consistency of the genetic signal across cohorts by splitting the 46 EUR-like discovery cohorts into two balanced halves, conducting IVW meta-analyses separately in each half and using LDSC to test the concordance of genetic signal across split halves. We identified novel lead SNPs in the current GWAS relative to past GWAS as those that were outside the same 250 kb locus as or LD-independent of a significant SNP in past GWAS of the Big Five. Finally, we quantified similarity in genetic signal among the Big Five by identifying overlapping 250 kb genetic loci across pairs of traits and by estimating genetic correlations between pairs of Big Five traits using LDSC.
Characterizing SNP heritability
We estimated h2 SNP for each personality trait using LDSC. As LDSC requires ancestrally homogenous data, we applied this method separately to GWAS summary statistics of participants with EUR-like genomes, AFR-like genomes and, for comparison purposes, EUR-like genomes in the subset of 10 cohorts that ascertained participants with both EUR-like and AFR-like genomes. We estimated polygenicity using stratified LD fourth moments regression68, and we estimated within-trait transancestry genetic correlations using POPCORN69.
We additionally applied LDSC to GWAS data on each individual cohort and then meta-analysed h2 SNP estimates across cohorts using fixed-effects and random-effects meta-analyses, with the R package metafor70. To quantify the effect of measurement error of the personality measure on h2 SNP , we conducted meta-regressions, where h2 SNP in a cohort was regressed on the Cronbach’s α reliability of its personality measure (reliabilities are shown in Supplementary Table 1). For each trait, we also used LDSC to correlate genetic effects across each pair of cohorts with positive h2 SNP estimates (1,781 total pairings) and used fixed-effects meta-analysis to obtain an average estimate for this cross-cohort genetic correlation.
... continue reading