ChEMBL34 mined for 191,732 ligand pairs to build a random baseline for optimization
Researchers filtered the PostgreSQL version of the CHEMBL34 database to find pairs of compounds tested against the same protein target, assay, activity type and publication, where one molecule was a simple chemical modification of the other (such as swapping a hydrogen for fluorine or a carbon for nitrogen). This produced 191,732 parent-analogue pairs, which the team supplemented by synthesizing new analogues from 18 selected parent compounds using RDKit-driven enumeration of feasible single-atom substitutions, each capped at $400 per 10 mg.