Rapid Hit Expansion Using CHEESE Search
From single molecules into millions of relevant compounds in a matter of seconds
Directed and precise chemical space expansion based on in vitro-obtained results is a powerful tool enabling further exploration of vast spaces of molecules sharing crucial features such as molecule shape or distribution of electrostatic charge. However, as open-source platforms and even commercial vendors offer their multi-billion chemical libraries, such combined chemical spaces represent an extremely rich dataset, typically lacking uniformity and means to explore molecules in a short, user-friendly time frame.
CHEESE Search offers a unique platform addressing such issues as it combines enumerated chemical spaces of the largest small-molecule databases available. It not only provides a vast ecosystem of a myriad ready-to-purchase molecules, but its key feature is its speed. Using the technology of CHEESE Embeddings allows users to search for millions of shape- or electrostatically similar molecules within single seconds with very low computational cost.
In this use case, we demonstrate a way to efficiently use CHEESE Search for rapid hit expansion, enabling an ultra-fast and cost-efficient exploration of chemical spaces in order to discover novel, structurally diverse, and highly relevant binders of the target protein within seconds.
For the use case involving a rapid hit expansion using the CHEESE Search platform, we selected BRD4 protein, which represents a challenging target both from the perspective of druggability and selectivity, since BRD4 bears two structurally very similar binding sites - BD1 and BD2. BRD4-BD2 is responsible for inflammatory and regulatory processes in the body, while BRD4-BD1 is a popular target for its regulation of cellular proliferation, and therefore has a great impact on the growth of cancer cells. Most known inhibitors of BRD4 are pan-BRD4, meaning they lack selectivity towards only one of the binding pockets.[1] However, the literature describes a limited set of BRD4-BD1-selective molecules. These molecules represent the basis of this study.
Starting with 22 in vitro verified hits from the literature[2], this relatively small dataset was expanded, using CHEESE Search, into 4.5 million electrostatically similar molecules.
CHEESE Search pre-sets were as follows:
All available databases
70% electrostatic similarity and higher
Accuracy - fast
Lipinski rule of 5 (Mw 300 - 550 Da)
The CHEESE Search-expanded dataset was processed using RDKit-based pipeline (accessible on our GitHub, see also figure 1), employing the following workflow:
SMILES validity check
Desalting
SMILES canonization
Deduplication
Synthetic accessibility scoring (SA score < 4.0)
PAINS filter (PAINS-A/B/C rule sets)
Murcko-Bemis scaffold clustering (one representative per unique generic Murcko scaffold)
MaxMin diversity selection (Morgan fingerprints, radius 2, 2048 bits; greedy Tanimoto clustering at cutoff 0.70, followed by property-biased MaxMin selection at Tanimoto ≤ 0.85)
The following step consisted of a docking study using third-party software. Both BRD4-BD1 and BRD4-BD2 were subjected to the docking experiment, as the goal was to discover compounds selective towards BRD4-BD1 only. The docking stage resulted in 400 relevant molecules. This set of candidate binders was used for a back-enrichment step, where compounds with Tanimoto similarity of 0.7 or higher to the 400 selected candidates were extracted from the initial 4.5 million-membered dataset. This step is crucial for reverse-extraction of relevant compounds which were omitted in the first clustering step. Such an enriched set of molecules was then subjected to the docking study again.
A thorough selection of candidates resulted in 12 novel molecules representing 11 unique scaffolds. Mean Tanimoto similarity of the candidate molecules to their respective nearest neighbor was 0.14, which signifies a diverse set of molecules (see figure 2). However, physicochemical properties were retained, as candidate molecules exhibit very similar key features and occupy very much the same chemical space as literature-based compounds (see figure 3).


This use case shows practical application of the CHEESE Search platform for the fast and cost-efficient expansion of limited number of in vitro validated hits into a large chemical space of relevant drug-like candidates exhibiting in silico activity towards the target, maintaining desired physicochemical properties of the lead compounds, while introducing rich structural novelty to the dataset.
Although this project has been done using CHEESE Search user interface, the platform itself offers many more functions. We aim for high customisability of our CHEESE Search platform. For larger-scale or automated workflows, a better accessibility can be achieved using the library of API requests.


