Results of CACHE Challenge #5
CACHE#5 participants who asked to be de-anonymized:
- Javier Vázquez, Brian Medel, Albert Herrero, Daniel Alvarez, Enric Herrero (Pharmacelera)
- Olexandr Isayev, Filipp Gusev, Ben Koby, Muhammed Aktolun, Polina Avdiunina, Cansin Ayvaz, Mudit Jaju, Hatice Gokcan & Maria Kurnikova (Carnegie Mellon University)
- Mykola Protopopov, Olga Tarkhanova, Olha Semenenko, Anna Kapeliukha, Oleksii Hrabovskyi & Oleksander Mosia (Chemspace)
- Andrea Volkamer, Antoine Lacour, Roman Joeres, MAx Rausch-Dupont, Alejandro Martínez León (Saarland University),& Joschka Groß (German Research Center for Artificial Intelligence).
- Sandro Cosconati, Benito Natale & Michele Roggia (DiSTABiF,University of Campania)
- Alexander Tropsha, Kushal Koirala (University of North Carolina Chapel Hill), Ryota Ashizawa , Stan Xiaogang Li, Dima Kozakov (University of Texas Austin) & Sergei Kotelnikov (Massachusetts Institute of Technology).
- Keunwan Park,Charuvaka Muvva, Young-Joon Ko & Dohyeon Kim (Korea Institute of Science and Technology)
- Andrea Paquola, Zhaoqian Su, Xiaohan Kuang, Hexuan Fan, Xuan Vu Nguyen, Yunchao Lance Liu & Wei Qiangqiang ( Alma AI, Inc.)
Receptor binding profiles, and agonist/antagonist functional data were generously provided by the National Institute of Mental Health's Psychoactive Drug Screening Program, Contract # 75N95023C00021 (NIMH PDSP). The NIMH PDSP is Directed by Bryan L. Roth MD, PhD at the University of North Carolina at Chapel Hill and Project Officer Jamie Driscoll at NIMH, Bethesda MD, USA. Drs. Magda Szewczyk and Dalia Barsyte-Lovejoy at the Structural Genomics Consortium ran antagonist assays on Round 2 compounds. Our thanks to Dr. Evert Homan at the Karolinska Institute for his preliminary analysis of the structural and chemical coverage of GPCRs and to Dr. Alex Tropsha at UNC who provided the list of MCHR1 ligands from the patent literature. Operational coordination: Dakota Treleaven, Yuchen Zhou at Conscience. Scientific oversight: Matthieu Schapira, Structural Genomics Consortium.
SUMMARY
Twenty-four computational teams predicted up to 100 MCHR1 antagonists each. The resulting 1455 compounds were tested experimentally in a competition assay and a functional assay. Forty-four compounds of interest from 16 teams were advanced to the hit expansion round (Round 2) where participants each selected up to 50 close analogs. The resulting 576 compounds were tested in the same 2 assays and a Hit Evaluation Committee scored each Round 1 compounds based on their activity, attractiveness for medicinal chemistry, and the activity of their Round 2 analogs. Nine groups were successful in predicting convincing MCHR1 ligands. While this was meant to be a ligand-based challenge (many known ligands, no close homolog in the PDB when predictions were made), most participants adopted structure-based approaches.
INTRODUCTION
- For CACHE #5, 24 participants from 12 countries used their computational workflows to predict small molecule antagonists of MCHR1, a GPCR involved in metabolic disorders. At the start of this challenge, over 3000 chemically diverse ligands had been reported in ChEMBL and the patent literature, and the structure of neither the target, nor close homologs were in the PDB. The challenge was to design chemically novel ligands.
- CACHE Challenges begin with a hit finding round (Round 1), where participants can nominate up to 100 commercially available compounds. Round 1 predictions are then experimentally tested and the resulting data returned to participants. In the hit-expansion round (Round 2), participants whose nominated compounds show some sign of activity in experimental binding assays (compounds of interest) can select up to 50 follow-up compounds. Together, these two rounds are designed to avoid both false positives and false negatives, which is critical to evaluate computational methods properly.
- In a separate exercise, computational teams are asked to predict hits from a library composed of the merged Round 1 selections from all participants. Here, all teams are screening the same library, and all compounds are experimentally characterized as active or inactive, but experimental data is blinded to participants.
THE CHALLENGE
- CACHE participants were asked to predict MCHR1 antagonists.
- After a double-blind peer review where each applicant reviewed 5 applications, 24 participants joined and completed the challenge, representing a diverse array of physics-based and ML methods. In spite of the absence of relevant structures in the PDB, most participants adopted structure-based approaches
Links to computational methods are provided in the table below
ROUND 1 RESULTS
All Round 1 experimental data are provided here and here
- 1455 compounds were first tested for their ability to compete with an MCHR1 antagonist (SNAP94847) using cell membrane preparations.
- 147 compounds showed at least 50% inhibition at 30 µM and were advanced to dose response experiment [concentration range: 10nM to 100 µM]
- A full dose response was observed for 26 compounds, leading to inhibitory activity (Ki) ranging from 170 nM to 30 µM while another 18 compounds had a partial allosteric modulator (PAM) profile.
- These 44 compounds were advanced to Round 2, even though some (mostly PAMs) showed signs of insolubility.
The 44 compounds were expected to behave as antagonists and were tested in functional GloSensor cAMP antagonist and agonist assays using transiently transfected HEK293 T cells. Thirteen compounds displayed weak or very weak antagonist activity. Some of the other compounds clearly interfered in both agonist and antagonist assays.
ROUND 2 RESULTS
All Round 2 data are available here ,here and here
- 576 compounds from 15 participants were tested in the SNAP94847 competition assay.
- 75 analogs of 16 Round 1 hits from 12 participants had a pKi > 5.
- 32 compounds - at least one compound from each chemical series - were tested in an antagonist assay and 18 from 7 participants were confirmed antagonists.
EVALUATION OF COMPUTATIONAL METHODS
- The biophysical data and SAR of Round 1 hits and their Round 2 analogs were blindly evaluated by an independent Hit Evaluation Committee composed of industry experts in medicinal chemistry (Hartmut Beck, Bayer & Lars Wortmann, Boehringer Ingelheim), computational
chemistry (Pat Walters, Relay Therapeutics ), biophysics (Anders Gunnarsson, AstraZeneca), and GPCR biophysics (XP Huang and Bryan Roth PDSP, University of North Carolina), leading to a final score assigned to each Round 1 hit - In addition to the experimental assay readouts, the Committee evaluated any chemical liability and chemical novelty compared with previously disclosed compounds. The annotated scores from the Hit Evaluation Committee are available here
- A first metric used to evaluate computational workflows is the best score obtained for any predicted hit, reflecting the efficiency of the workflow in delivering a good chemical starting point for follow-up medicinal chemistry.
- A second metric is the aggregated score (i.e. sum of all scores) across all confirmed hits predicted by a workflow which better reflects hit rate.
- Each participant was provided with the unlabelled list of all 1455 compounds predicted in Round 1 and asked to predict up to 100 actives from this list. A third metric is the aggregated score of compounds predicted active from this list
- The results are provided in the table below.
- Finally, a plot highlighting the potency and chemical novelty of Round 2 analogs is shown below, and reflects understanding of the chemistry driving protein binding (and commercial availability of analogs).
Nine workflows convincingly predicted experimentally confirmed MCHR1 ligands. Two of them were purely ligand-based (WF2127 and WF2123), others were hybrid or purely receptor based.
Table 1: Scoring CACHE #5 workflows. Each compound was assigned a score by a panel of medicinal chemistry and GPCR biophysics experts (the IDs of participants were blinded). The score of the best predicted compound, the sum of scores for all compounds predicted from Enamine, and the sum of scores for all compound predicted active from the merged selection of 1455 Round 1 compounds are provided (N/A: no prediction submitted).
| Workflow | Participant | Top Scores | Aggregated Scores | Normalized Scores Merged Selection | Hit Finding Strategy |
| 2127 | Javier Vázquez [Pharmacelera] | 23.5 | 42.0 | 89.9 | Ligand-based fragment matching fragments to Enamine building blocks, incorporating quantum chemical descriptors |
| 2132 | Olexandr Isayev [CMU] | 29.0 | 44.8 | N/A | Ligand-based QSAR -> Docking + refinement with QM-trained ML force fields |
| 2156 | Alex Powers [ Stanford U] | 17.3 | 33.3 | N/A | Receptor-based generative model tuned on Enamine REAL |
| 2150 | Mykola Protopopov [Chemspace] | 16.9 | 16.9 | 40.2 | V-SYNTHES– combinatorial enumeration of best-docked fragments + active learning screening of enumerated library and mmPBSA/GBSA calculations |
| 2153 | Andrea Volkamer [UdS] | 16.3 | 16.3 | 62.0 | Ensemble of diverse DL methods including post processing via docking and FEP |
| 2123 | Sandro Cosconati [U.Campania] | 16.0 | 16.0 | 33.9 | ML-enforced ligand-based virtual screening with PyRMD |
| 2141 | 14.8 | 14.8 | 0.0 | Docking-based active learning + pharmacophore filtering | |
| 2144 | Alexander Tropsha [UNC] & Dima Kozakov [UT Austin] | 14.6 | 19.8 | N/A | SBDD workflow only, QSAR-based virtual screening only, and consensus hits from both workflows |
| 2151 | Keunwan Park [KIST] | 14.2 | 14.2 | 49.2 | Evolutionary chemical binding similarity model + ML-based ranking and filtering |
| 2128 | 9.8 | 14.0 | N/A | ||
| 2140 | 9.0 | 10.3 | 35.8 | ||
| 2142 | Andrea Paquola [Alma AI, Inc] | 8.3 | 14.1 | 34.9 | |
| 2133 | 2.5 | 2.5 | 53.4 | ||
| 2122 | 0.0 | 0.0 | N/A | ||
| 2145 | 0.0 | 0.0 | N/A | ||
| 2143 | 0.0 | 0.0 | N/A | ||
| 2146 | 0.0 | 0.0 | 48.4 | ||
| 2154 | 0.0 | 0.0 | 45.3 | ||
| 2158 | 0.0 | 0.0 | N/A | ||
| 2325 | 0.0 | 0.0 | 40.2 | ||
| 2157 | 0.0 | 0.0 | 0.0 | ||
| 2159 | 0.0 | 0.0 | 12.7 |

Figure 1: Potency and novelty of Round 2 analogs. pKi are derived from competition experiments. Tanimoto distances were calculated with ICM (Molsoft). Workflows with less potent/novel analogs in Round #2 may rank better in Table 1 if the antagonist activity or medicinal chemistry tractability of parent molecules was judged superior.