Plasmids and inserts
Sequences and accompanying information are given in Supplementary Tables 2–4. In brief, we selected Codebook TFs (and their DBDs) from information published in a previous study1 and posted at https://humantfs.ccbr.utoronto.ca. Inserts named with a ‘-FL’ suffix correspond to the full-length ORF of a representative isoform of the protein. Those with a ‘-DBD’ suffix contain all of the predicted DBDs in the protein flanked by either 50 amino acids or up to the amino or carboxy terminus of the protein. Those with a ‘-DBD1’, ‘-DBD2’ or ‘-DBD3’ suffix contain a subset of the DBDs present in the proteins; these were manually designed, mainly for large C2H2-zf arrays. Inserts were obtained as recoded synthetic ORFs (BioBasic) flanked by AscI and SbfI sites and subcloned into up to three plasmids: (1) pTH13195, a tetracycline-inducible, N-terminal eGFP-tagged expression vector with FLiP-in recombinase sites8; (2) pTH6838, a T7-promoter driven, N-terminal GST-tagged bacterial expression vector53; and (3) pTH16500 (pF3A-ResEnz-egfp), a SP6-promoter driven, N-terminal eGFP-tagged bacterial expression vector, modified from pF3A–eGFP7 to contain the two restriction sites after eGFP.
Protein production
Each experiment used a protein expressed from one of the following systems: (1) FLiP-in HEK293 cells (Thermo Fisher Scientific, R78007), induced with doxycycline for 24 h, used for inserts in pTH13195; (2) PURExpress T7 recombinant IVT system (NEB, E6800L), for inserts in pTH6838; or (3) SP6-driven wheat germ extract-based IVT (Promega, L3260), for inserts in pTH16500.
DNA-binding assays
We followed previously described protocols for ChIP–seq8, PBMs50 and SMiLE-seq7. Detailed descriptions of GHT-SELEX, HT-SELEX, ChIP–seq and SMiLE-seq data collection and initial analyses are provided in the accompanying papers10,11,12,13. For PBMs, we analysed proteins on two different universal PBM arrays (HK and ME), with differing probe sequences54, but did not analyse most C2H2-zfs owing to low success rates with this assay, presumably due to long binding sites. By default, we analysed each protein twice by ChIP–seq (with full-length constructs only)11. We analysed each construct by HT-SELEX and GHT-SELEX once by default (that is, as full-length and DBDs), and some constructs were analysed multiple times to examine the impact of experimental variables as the GHT-SELEX method was developed10. We ran SMiLE-seq assays for 278 TFs, which corresponded to 388 constructs. Controls were omitted given that most were selected because they already had published SMiLE-seq data, which resulted in 299 TFs with SMiLE-seq-derived motifs. A subset of Codebook proteins (mainly those with unknown DBDs) were also omitted owing to lack of success in all other assays. A randomly selected subset of 82 constructs was analysed multiple times to assess reproducibility across SMiLE-seq experiments. Anti-eGFP antibody (Ab290, Abcam) was used as the major antibody across all the assays (ChIP–seq, GHT-SELEX and western blots), with method-specific amounts10,11. Each ChIP reaction used 2 µl undiluted polyclonal antiserum, corresponding to 10 µg total IgG, immobilized on 60 µl protein G magnetic bead suspension (Dynabead 10004D, ThermoFisher). For HT-SELEX and GHT-SELEX, antibody-bead master mixes were prepared by immobilizing 6 µl antiserum on 100 µl protein G sepharose bead slurry (Cytiva, 28-9670-70), and 1 µl aliquots of the resulting mixture were used for each selection, which corresponded to approximately 0.06 µl antiserum (300 ng total IgG) per reaction. For western blotting, membranes were incubated in 15 ml antibody solution diluted 1:5,000, which corresponded to approximately 15 µg total IgG per membrane.
Data processing and motif derivation
The accompanying paper12 describes motif derivation and evaluation in detail. In brief, after initial preprocessing, we obtained a set of ‘true positive’ (likely to be bound) sequences for each individual experiment. A total of 721 out of 4,873 experiments were removed at this step owing to a low number of likely bound sequences or to other technical issues, as documented in Supplementary Table 4. We then applied a suite of tools (listed in Supplementary Table 15) to a training subset of the data from each experiment and tested the resulting motifs on a test subset of the data from the same experiment and on the independent data for the same TF (that is, the test sets from all other experiments performed for the same TF). We used a binary classification regimen for all experiments and all motifs and scored the motifs using a variety of criteria, including the area under the receiver operating characteristic curve (AUROC) and the area under the precision–recall curve. The full set of motifs (as PWMs) is available at Zenodo55, and an interactive browser is available at https://mex.autosome.org.
Systematic filtering of artefactual motifs
While curating the datasets and assembling a reliable motif collection, we accounted for enrichment of similar artefact motifs. These recurrent DNA patterns were detected owing to systematic experimental noise or peculiarities of particular motif discovery tools. For example, in HT-SELEX experiments, an ACGACG motif was often enriched. This sequence is a presumed artefact as it matches the constant flanking region. In experiments with cell lysates, motifs of abundant native proteins in HEK293 cells were sometimes enriched (for example, NFI, YY1 and ETS-family TFs). To minimize the influence of these artefacts, we performed the following tasks: (1) manually compiled a list of recurrent artefact motifs (Supplementary Table 16); and (2) scanned the entire motif collection with MACRO-APE56 and removed motifs that were highly similar to those in our catalogue of artefacts. We also removed motifs that matched the constant, non-variable regions of the DNA used in HT-SELEX, GHT-SELEX and SMiLE-seq experiments. We did not filter out ETS-related motifs for ETS-family positive controls, such as ELF3, FLI1 and GABPA. Subsequent to expert curation, we confirmed that enriched k-mers in HT-SELEX experiments did not correspond to potential artefacts associated with individual expression systems10.
... continue reading