Biospecimen processing and quality control
Before extracting nucleic acids, pathology quality control (QC) was performed on all frozen and FFPE tumours and normal specimens. From frozen tissue, a 30 mg or smaller piece was prepared, or an equivalent amount of scrolls were cut from a FFPE block, with sections stained with haematoxylin and eosin taken from the top and bottom of each specimen. The stained slides were scanned and their pathology reviewed to confirm that the tumour was consistent with the reported histology, and to assess the per cent tumour nuclei, per cent necrosis and other pathological features. The tumour nucleus and necrosis percentages from the top and bottom slides were averaged, and tumour specimens with an average of ≥50% tumour nuclei and ≤20% necrosis were submitted for nucleic acid extraction. Normal tissues were rejected if they had any detectable tumour cells.
DNA was extracted from normal blood and saliva specimens. DNA and RNA were co-extracted from tumours, solid normal tissues and cancer models. Before extraction, cancer models were first washed with ice-cold PBS to remove residual Matrigel. Frozen tissues and cancer models were homogenized using a Qiagen TissueLyser, and RNA and DNA were extracted using a modification of the DNA/RNA AllPrep kit (Qiagen). The homogenate was applied to a Qiagen DNA column, and the flow-through was processed using a mirVana miRNA Isolation kit (Ambion). FFPE specimens were deparaffinized and cells were lysed. The pellet underwent DNA extraction using an AllPrep FFPE kit (Qiagen), whereas the supernatant underwent RNA extraction using a Highpure miRNA kit (Roche). Blood specimens were extracted using a QiaAmp DNA Blood Midi kit (Qiagen), and saliva was extracted using a Gentra Puregene Buccal Cell kit (Qiagen).
DNA was quantified by PicoGreen assay and RNA was quantified by measuring the absorbance at 260 nm with a UV spectrophotometer. DNA quality was assessed by 1% agarose gel electrophoresis to confirm high-molecular-weight fragments. RNA was analysed using a RNA6000 Nano assay (Agilent) on an Agilent Bioanalyzer, which returns an RNA integrity number (RIN) for RNA from frozen tissues, or a DV200 for RNA extracted from FFPE specimens. All extracted DNA was subjected to a custom Sequenom single-nucleotide polymorphism (SNP) panel or an AmpFISTR Identifiler (Applied Biosystems) to verify that all specimens representing a case were derived from the same patient.
Genome sequencing
Data processing
The Genomics Data Commons (GDC) DNA-seq alignment pipeline59, was used to map WGS reads to a customized version of GRCh38. In brief, reads were aligned using BWA-MEM (v.0.7.15)60, followed by sorting, merging and duplicate-marking using Picard Tools (v2.26.10). Base quality scores were recalibrated using GATK (v.3.7.0)58 and BQSR. Sample contamination levels were estimated using GATK ContEst, and samples with contamination levels exceeding 4% were excluded from further analyses.
SNV and indel calling
The New York Genome Center pipeline. The New York Genome Center (NYGC; v.6) somatic SNV–indel-calling pipeline61 was run for each tumour–normal and model–normal pair, starting from the GDC-aligned BAM files. In brief, SNVs, multi-nucleotide variants (MNVs) and indels were called using MuTect2 (GATK v.4.0.5.1)62, Strelka2 (v.2.9.3)63 and Lancet (v.1.0.7)64. Indels were also called using SvABA (v.0.2.12)65. Candidate indels predicted by Manta (v.1.4.0)66 were used as input to Strelka2, as per developer recommendations. Variants were merged across callers and annotated using Ensembl (v.93)67, COSMIC (v.86)47, 1000Genomes (Phase3)68, ClinVar (201706)69, PolyPhen (v.2.2.2)70, SIFT (v.5.2.2)71, FATHMM (v.2.1)72, gnomAD (r.2.0.1)73 and dbSNP (v.150)74, using Variant Effect Predictor (v.93.2)75. From the final callset of SNVs and indels, we filtered those that met the following criteria: occurred in two or more individuals in a panel of normal samples61; those that had a minor allele frequency (MAF) of at least 1% in 1000Genomes or in gnomAD; had a tumour VAF less than 0.0001; had a normal VAF greater than 0.2; had depth less than 2 in either the tumour or the normal sample; or a VAF in the normal sample greater than in the tumour sample. The aforementioned panel of normal samples was constructed from 242 unrelated individuals, of which 148 were sequenced using HiSeqX in the Illumina Polaris project76, 73 were sequenced using HiSeqX at the NYGC, 11 were sequenced using NovaSeq at the NYGC, and 10 were sequenced on both HiSeqX and NovaSeq platforms at NYGC.
Broad pipeline (whole-exome sequencing).Whole-exome sequencing (WES) BAM files were processed as described in the section ‘Data processing’. Variant calling was performed as previously described77. SNVs were called using MuTect and small indels were called using Strelka. Tumour-in-normal contamination estimation was performed using deTiN, and cross-participant contamination was estimated using ContEst. WES variant calling was performed as previously described77. A similar pipeline was used for processing variant calls from WGS.
... continue reading