Ethics
The work reported on TwinsUK samples was carried out with informed consent under TwinsUK BioBank ethics, approved by the North West–Liverpool Central Research Ethics Committee (REC reference 19/NW/0187), IRAS ID 258513 and earlier approvals granted to TwinsUK by the St Thomas’ Hospital Research Ethics Committee, later the London–Westminster Research Ethics Committee (REC reference EC04/015). SL samples were obtained from Discovery Life Sciences as discarded medical waste in accordance with IRB protocol DLS-BB050. These were collected as part of standard-of-care testing ordered by the patient’s physician and excess material released for research.
Sperm sample extraction and sequencing
We obtained 15 sperm samples from 13 donors of European ancestry. Eleven individuals donated one sample, and two individuals donated two samples at different ages (AA1-s1 and AA1-s2; AN-s1 and AN-s2), with ages at donation varying between 24 and 74 (Supplementary Table 1). Samples AA2-t1 and AA2-t2 are monozygotic twins. The TwinsUK cohort consists of more than 14,000 volunteers, mostly middle-aged female individuals, who have participated over the past 30 years in a longitudinal cohort study. This has included lifestyle and health questionnaires, biomedical measurements, biological sample collection and the generation of multi-omics profiles (such as genetics, metagenomics and metabolomics), over multiple visits.
For the TwinsUK samples, we obtained high molecular weight (HMW) DNA from bulk sperm samples using the Circulomics Nanobind Tissue Kit (102-302-100) and NEB Monarch HMW DNA Extraction Kit for Cells & Blood (T3050) UHMW with some modifications to account for tighter packing of the sperm genome. We supplemented the lysis solution with 150 mM 1,4-dithiothreitol to disrupt protamine disulfide bonds and extended proteinase K incubation from 30 min to 2 h. The rest of the Circulomics HMW DNA extraction protocol remains unchanged. HMW DNA was sheared to 10–14 kb DNA fragments using Megaruptor 3 system (B06010003) with speed setting 30. CCS sequencing libraries were constructed according to the standard CCS library preparation protocol 1.0 (100-222−300), and the libraries were sequenced using a Sequel IIe and Revio instrument at the Wellcome Sanger Institute with full-resolution base quality scores.
For the SL data, samples were thawed on ice and cells were pelleted in an isotonic sperm wash solution to remove debris and reduce somatic cell contamination. DNA was extracted using the NEB Monarch HMW DNA Extraction Kit for Cells & Blood (T3050) UHMW protocol, with some modifications to the cell lysis step. Specifically, cells were digested at 56 °C for 1 h at 300 rpm with 100 µl Nuclei Prep Buffer, 100 µl Nuclei Lysis Buffer, 10 µl Proteinase K (NEB P8107, 20 mg ml−1), and 10 µl 1 M 1,4-dithiothreitol (GoldBio in dH 2 O, final concentration ~50 mM). An additional 20-minute digestion was then performed with 5 µl RNase A (NEB T3018, 20 mg ml−1). The rest of the protocol remained unchanged. HudsonAlpha performed HMW DNA quality control, CCS library preparation, and sequencing on the PacBio Revio sequencing instrument.
Blood sample sequencing data analysis
We obtained raw PacBio CCS sequencing data and assemblies from the Platinum Pedigree dataset. Flow cells containing multiple samples were excluded due to cross-sample contamination caused by demultiplexing errors. Additionally, we excluded the four first-generation samples, as they were derived from cell lines. This resulted in a final dataset of 12 samples.
De novo haplotype-resolved assembly
We used hifiasm30 (v0.19.5-r592, default parameters, HiFi only mode) to construct a haplotype-resolved de novo assembly for each sample. The hifiasm assembler outputs a partially phased assembly graph for the two haplotypes, which we converted into two fasta sequence files per sample.
... continue reading