Abstract
Chemical and biological data are the cornerstone of modern drug discovery programs. Finding qualitative yet better quantitative relationships between chemical structures and biological activity has been long pursued in medicinal chemistry and drug discovery. With the rapid increase and deployment of the predictive machine and deep learning methods, as well as the renewed interest in the de novo design of compound libraries to enlarge the medicinally relevant chemical space, the balance between quantity and quality of data are becoming a central point in the discussion of the type of data sets needed. Although there is a general notion that the more data, the better, it is also true that its quality is crucial despite the size of the data itself. Furthermore, the active versus inactive compounds ratio balance is also a major consideration. This review discusses the most common public data sets currently used as benchmarks to develop predictive and classification models used in de novo design. We point out the need to continue disclosing inactive compounds and negative data in peer-reviewed publications and public repositories and promote the balance between the positive (Yang) and negative (Yin) bioactivity data. We emphasize the importance of reconsidering drug discovery initiatives regarding both the utilization and classification of data.
1 Introduction
Data and the increasing role of predictive models, including machine and deep learning (; ), are the cornerstone of modern drug discovery programs (Zhang et al., 2022). The increasing use of computational methods that recently included deep learning is reducing the time and financial costs of finding drug candidates (Zhang et al., 2022). For instance, computer-aided drug design (CADD) has led to the discovery of more than seventy approved drugs (Sabe et al., 2021) including remdesivir as an emergency treatment against SARS-CoV-2 in 2021 ().
CADD methods are typically divided into two main categories, structure-based drug design (SBDD) and ligand-based drug design (LBDD) that rely on the three-dimensional (3D) structure data available for one or more molecular targets, or the structure-activity data of ligands, respectively. Examples of deep learning applications in SBDD include AlphaFold to assist in homology modeling, and DiffDock in molecular docking. AlphaFold predicts 3D protein structures according to their amino acid sequences (), and DiffDock predicts the binding mode between the ligand and specific protein target (). One of the most notable approaches in LBDD are quantitative structure-activity relationships (QSAR) (). Current QSAR methods use machine learning and deep learning (Soares et al., 2022) that can be divided into linear methods and nonlinear methods (Patel et al., 2014; ). Linear methods include linear regression, multiple linear regression, partial least squares, and principal component analysis (Patel et al., 2014). Nonlinear methods include artificial neural networks, k-nearest neighbors, and Bayesian neural nets, to name a few examples (Patel et al., 2014; ).
Advances in deep learning models have a significant progress in molecule generation, representing a big step forward in bridging the gap between chemical entities and drug-like properties (). Deep learning algorithms are currently used in the renewed interest in the de novo design of chemical libraries. In 2020, the successful application of deep learning in drug discovery, that included the de novo design using deep learning, was selected by the Massachusetts Institute of Technology Technology Review as one of the top ten breakthrough technologies ().
De novo design is aimed at generating new chemical entities (NCE) with desired properties (Palazzesi and Pozzan, 2022). De novo design based on deep learning algorithms (Palazzesi and Pozzan, 2022) requires a large number of compounds that may demand significant computational resources. However, bioactivity data for a biological endpoint is not always sufficient. The lack of data has led to the development of new methods for compound selection and applications for deep learning algorithms are being developed ().
Knowledge-based drug design frequently involves quality data (Perron et al., 2022b) to develop models with useful predictions (Schneider et al., 2020). To this end, rethinking the methodologies used for drug discovery and development campaigns is crucial. The quality of data sets, decoy data sets and inactive compounds used in predictive models, and de novo design models need to be reviewed and discussed.
The main purpose of this manuscript is discussing the importance of quality data, decoy data sets, and the balance needed between inactive (i.e., “Yin”) and active (“Yang”) compounds currently employed in de novo design and developing predictive models of biological activity to generate NCE. Following up on previous studies (Schneider et al., 2020; ; ), we comment on the need to rethink the way to drug design and develop campaigns. The manuscript is organized into four main sections. After this Introduction, Section 2 presents an overview of de novo design. Section 3 discusses the main public data sources used to develop predictive models. Section 4 discusses criteria to generate quality data sets. The last section presents a summary of conclusions and perspectives.
2 De novo design overview
De novo design aims to generate new chemical structures from scratch with desired predicted properties, e.g., absorption, distribution, metabolism, excretion, toxicity (ADMET), other drug-likeness properties, and biological activities (Palazzesi and Pozzan, 2022). The two main strategies for de novo design can be classified into SBDD and LBDD (vide supra) (Zhang et al., 2022). A recent example of a structured-based de novo design is the RELATION model that learns from the desired geometric features of protein-ligand complexes to generate new molecules (Wang et al., 2022). The generation process applies a fragment-based strategy given an initial chemical scaffold embedded in the binding site of the target protein. The pre-trained model generates molecules iteratively by sequentially adding, deleting, inserting, or replacing and linking fragments (Zhang et al., 2022).
In contrast, ligand-oriented de novo design focuses on the ligands themselves, thereby generating compounds with new chemical structures with novel scaffolds from active compounds while optimizing the desired properties (Xie et al., 2022). A general workflow is schematically summarized in Figure 1 which has seven main steps (; Zhang et al., 2022): 1) Selecting compound data sets from public or in-house sources (further discussed in Section 3); 2) Filtering molecular data sets with desired properties such as drug-likeness. In the example of Figure 1 a data set with three subsets of compounds is represented with a star, triangle, and circle, respectively. The compounds represented with a star have drug-like properties (; Veber et al., 2002); those represented with triangles comply with some of the drug-likeness properties, and those represented with circles are not compliant. Other approaches to select compounds from the data sets use molecular fingerprints () or filter compounds directly via similarity-based virtual screening instead of designing NCE from scratch (Tong et al., 2021). 3) Selecting the molecular representation as a basis to learn and represent the structures and properties of molecules, e.g., SMILES (Weininger, 1988), SELFIES () or molecular graphs (Simonovsky and Komodakis, 2018). 4) Developing and validating the model for molecule generation using metrics such as the operating characteristic curve. 5) Optimizing the model by combining reinforcement learning and property prediction (). 6) Generating molecules de novo, 7) Assessing the biological activity of the compounds designed in relevant in vitro or in vivo models.
FIGURE 1
Deep learning, currently used in ligand-based de novo design, learns the probability distribution of molecular data and generates continuous or discrete latent representations for molecules with property optimization (
Ligand´s properties can be optimized in two steps: 1) property-based generation, wherein models would learn the chemical space of molecules with desirable properties; and 2) novel molecules are generated within a desired property space (
Ligand-based de novo design using DNN (Palazzesi and Pozzan, 2022) requires a large number of compounds that demand more computational resources. The DNN architecture is prone to problems because of fitting numerous parameters. For this reason, a large training data set is needed to reduce the risk of overfitting. However, sufficient bioactivity data for a biological endpoint is not always available (Wu et al., 2018). The lack of sufficient data has led to using methods for compound selection or the development of new methods for compound selection. Altae-Tran et al. (
The availability of gold standard datasets as well as independently generated data sets are valuable in generating well-performing models (Vamathevan et al., 2019). Dissimilarity-based compound selection could be improved if one focused the selection on a structural diverse dataset (for instance derived from natural products). Some approaches proposed suggest using quality data sets using a dissimilarity-based compound selection method such as the MaxMin or MaxSum algorithms (Leach and Gilleteds, 2007). Recently, we reported the use of the MaxMin algorithm for the selection of natural product subsets (
3 Main sources of data sets used to develop generative and predictive models
3.1 Current status of reference and benchmark datasets
The first step in de novo design is to select, from the vast chemical space, the appropriate subset of all possible molecules for a desired biological activity (Schneider et al., 2000). To have an idea, the size of the chemical space has been estimated at around 1060 small molecules and between 1020–1024 for all molecules up to 30 atoms that comply with Lipinski’s rule-of-five (Reymond, 2015). According to Yang et al. compound data sets can be classified into on-demand databases, collections containing bioactivity data, compounds databases commercially available, and natural products databases (Yang et al., 2019). Herein, we include benchmark, decoy and inactive compounds data sets as others categories as illustrated in Figure 2. In this figure, on-demand databases are further divided into commercially available (e.g., Enamine-REAL, CHEMriya and Freedom Space) (
FIGURE 2

Classification of compound databases and representative examples of each one. For the discussion of this manuscript, databases are split into six main categories: on-demand, commercial availability, bioactivity, natural products, benchmark and decoys.
Among the different types of chemical databases, de novo design employs libraries from different categories outlined in Figure 2. Specific examples are ChEMBL (
TABLE 1
| Data sets | Category | Description | Ref. |
|---|---|---|---|
| ChEMBL | Bioactivity | Database with 2,354,965 bioactive drug-like small molecules with 2D structures and calculated properties. | |
| PubChem | Bioactivity | Database at the US National Institutes of Health with 115 million compounds. It includes names, molecular formulas, structures, physical properties, and biological activities. | |
| DrugBank | Bioactivity | Version 5.1.10 contains 15,448 drug entries including 2,740 approved small molecule drugs. | Wishart et al. (2006) |
| ZINC-22 | Commercial | Database with over 37 billion enumerated, searchable, commercially available compounds in 2D. | Tingle et al. (2023) |
| CHEMriya | On-demand | Database with 12 billion novel and synthetically feasible small molecules. | |
| Freedom Space (Chemspace) | On-demand | Database with 201 million molecules; 73% of its compounds comply with drug-likeness properties. | |
| Enamine-REAL | On-demand | Database with 6 billion synthetic compounds that comply with drug-likeness properties. | |
| MoleculeNet | Benchmark | Compilation of 17 datasets with over 700,000 compounds in total used for comparison of different machine learning algorithms. | Wu et al. (2018) |
| MOSES | Benchmark | Dataset with 1,936,962 molecules from ZINC Clean Lead suitable for hit identification and ADMET optimization. It does have metrics to detect common issues in generative models such as overfitting or if the model does not limit to producing only a few typical molecules. | Polykovskiy et al. (2020) |
Main sources of public molecular data sets used in de novo design.
3.2 On-demand databases
Early approaches to ligand-based de novo design involved fragment compounds into unique building blocks which could be recombined to make new molecules. A number of commercial suppliers of chemical samples offer large make-on-demand collections that can be reliably synthesized because the building blocks are available as well as the synthetic routes and methods (Warr et al., 2022;
3.3 Commercially available databases
One of the largest and long-standing compendiums of commercially available compounds in ZINC. The most recent version, ZINC-22 (Tingle et al., 2023) contains over 37 billion enumerated, searchable, commercially available compounds in 2D. Over 4.5 billion have been built in biologically relevant ready-to-dock 3D formats (Tingle et al., 2023). Some examples of de novo design using ZINC include the design of inhibitors of DDR1 (discoidin domain receptor 1, a kinase target implicated in fibrosis and other diseases) (Zhavoronkov et al., 2019) and compounds with activity towards the dopamine receptor D2 (
3.4 Bioactivity databases
De novo design based on deep learning algorithms frequently use PubChem, ChEMBL, and DrugBank to select subsets of compounds focused on a biological target or biological endpoint as the design of ligands (
3.5 Natural product databases
Natural product databases (
Privileged structures were defined by Evans et al. (
Representative natural product datasets that can be used in de novo design are Collection of Open NatUral ProdUcTs (COCONUT) (Sorokina et al., 2021), SuperNatural 3.0 (
TABLE 2
| Data sets | Description | Ref. |
|---|---|---|
| COCONUT | Extensive database with 406,076 unique structures. | Sorokina et al. (2021) |
| SuperNatural 3.0 | A database with 449 058 natural compounds and derivatives. It includes chemical structure, physicochemical information, information on pathways, mechanism of action, toxicity, vendor information if available, drug-like chemical space prediction for several diseases such as antiviral, antibacterial, antimalarial, anticancer, and target-specific cells. | |
| UNPD | Second-largest database with around 229,000 natural products that contain chirality information. | |
| TCM Database@Taiwan | Database with more than 20,000 pure compounds isolated from 453 TCM ingredients. | |
| IMPPAT | Database of 9,596 phytochemicals from 1,742 Indian medicinal plants. | |
| AfroDB | Compound collection with more than 1,000 compounds from African medicinal plants. | |
| NuBBEDB | Brazilian database with 2,223 natural products encoding as SMILES, InChI, and InChIKey strings, Ro5 and Veber descriptors, source, therapeutic effect, and reference. | Valli et al. (2013),Pilon et al. (2017),Saldívar-González et al. (2019) |
| SistematX | Brazilian database with 9,514 unique secondary metabolites encoding as SMILES, InChI, and InChIKey strings, and include physicochemical drug-like descriptors, predicted biological activities, and reference. | Scotti et al. (2018), |
| CIFPMA | Database developed at the University of Panama. It contains natural products that have been tested in over 25 in vitro and in vivo bioassays, for different therapeutic targets. | Olmedo et al. (2017),Olmedo and Medina-Franco (2020) |
| PeruNPDB | Peru database developed at the Catholic University of Santa Maria. The current version has 280 natural products from animals and plants. | |
| BIOFACQUIM | Mexican database with structures of 531 natural products isolated and characterized at UNAM and other Mexican institutions. | Pilón-Jiménez et al. (2019),Sánchez-Cruz et al. (2019) |
| UNIIQUIM | Mexican database with 1,112 plant natural products mostly isolated and characterized at the Institute of Chemistry of the UNAM. | UNIIQUIM (2015) |
Examples of natural product databases in the public domain.
Other libraries of natural products with an emphasis on commercial availability are listed on the NIH website (
SuperNatural 3.0, COCONUT and UNPD are the most extensive natural product databases. SuperNatural 3.0 (
Several public natural products databases compile the compounds isolated and characterized from a geographical region or the country of origin as China, India and Africa. For instance, Chinese Traditional Medicine (TCM) Database@Taiwan (
Representative Latin American databases (
3.6 Benchmark databases
The development of reliable machine learning algorithms has been limited due to the lack of standard benchmark datasets to compare the efficacy of the methods proposed (
3.7 Current decoy data sets and inactive compounds
Accuracy of predictive models depends on data quality and quantity. Also, the balance between active and inactive compounds is important, which remains an issue to resolve. Historically, the publication of active compounds in a given assay or with a particular endpoint has been prioritized over inactive molecules. For example, a recent comprehensive analysis of published screening bioactivity data shows that in ChEMBL V.29 (release in 2022) there is a large number of active compounds (ca. 71%) with respect to the inactive ones (ca. 31%); contrary to what it would be expected (
Decoy data sets have been developed in an attempt to reduce the gap between inactive (or negative) and active compounds. Decoy molecules are assumed non-active but have high physicochemical property similarity (but not topologically) to reference compounds (Réau et al., 2018). Decoys are useful to evaluate benchmark models that were assembled in the absence of inactive compounds experimentally measured (
TABLE 3
| Datasets with active and inactive compounds | Criteria to select inactive data | Ref. |
|---|---|---|
| ChEMBL | Reported activity data. | |
| PubChem | ||
| Binding DB | Reported ligand-receptor affinity. | |
| Decoy datasets | Common decoy selection criteria | |
| ZINC | Compounds that share drug-like properties with the reference (active) compounds. | Tingle et al. (2023) |
| DUD-E | ||
| DUD | Database with 2950 annotated ligands and 95,316 property-matched decoys for 40 targets. | |
| MUV | Compounds that share structural similarity with active reported compounds. | Rohrer and Baumann (2009) |
| DEKOIS 2.0 | Compounds that share drug-like properties and structural similarity with the reference (active) compounds. | |
| Decoy tools | Common decoy compound selection criteria | |
| DecoyFinder | Allows the automatic creation of datasets of compounds with physicochemical similarity and without structural similarity respect to the reference (active) compounds. | |
| RADER | Allows the automatic generation of datasets of compounds with physicochemical and structural similarity with respect to the reference (active) compounds. | Wang et al. (2017) |
| ZINC pharmer | Enables the automatic identification of compounds with pharmacophore similarity with respect to the reference (active and inactive) compounds. | |
| Decoy Developer | Allows the automatic generation of peptides decoys. | Shipman et al. (2019) |
Examples of potential inactive and decoy resources for enriching de novo design models.
Decoy compounds have been used to describe, explore, and expand the knowledge of active molecules. For example, rationalizing the physicochemical, chemical, biological, and clinical data of active compounds (
TABLE 4
| Approach | Purpose of using decoy sets | Ref. |
|---|---|---|
| Ligand-based | • Validation of new protocols and scoring functions based on similarity metrics and 3D shape. | |
| • Improvement of the accuracy of AI-based models. | ||
| • Improvement of the accuracy of QSAR models. | ||
| • Enrichment of inactive “dark regions” in chemical space. | ||
| Structure-based | • Validation of new protocols and scoring functions based in docking, molecular dynamics, and pharmacophore modeling. | |
| • Peptide and protein design. |
Examples of applications of decoys in de novo design.
4 Criteria to generate compound datasets with high quality
The quality of a data set is multifaceted. Commonly, it is associated with the experimental reproducibility of each data point and the experimental similarities between the protocols used to derive such data. Another important aspect of data quality is the balance between active and inactive compound. The latter is specially a challenge in public data sets due to the overall lack of published negative data. Finding qualitative yet better quantitative relationships between chemical structures and biological activity has been long pursued in medicinal chemistry and drug discovery. With the rapid increase and deployment of the predictive machine and deep learning methods, as well as the increased interest in the de novo design of chemical libraries (
TABLE 5
| Criteria | Brief description | Ref. |
|---|---|---|
| Balance | • Quality and quantity data allow the exploration of substantial regions of chemical space. | Scannell et al. (2022); Yang et al. (2023) |
| Quality (confidence) data | • The reliability of the activity data (active or inactive) is crucial to develop predictive models. This is the activity data reproducibility. | |
| Diversity | • Datasets with a high chemical and structural diversity improve the generation of novel molecules. | Saldívar-González and Medina-Franco (2022) |
| Preparation or curation | • Dataset curation must be focused on one or multiple drug targets. Therefore, molecular descriptors and the cut-off threshold used for the curated must be properly selected. • Dataset should be oriented to resolve specific outcomes and avoid Pan-Assay Interference Compounds (PAINS) structures or chemical structures related to side effects. • In small datasets it is very important to have as much accurate data as possible. The maximum observable accuracy of classification models also depends on the experimental uncertainty and the distribution of the measured values. For instance, datasets with large noise are not recommended for the comparison of different models. | |
| Complete information | • According to the main objective of each project, the dataset used must contain reliable data related to the project’s objective. For example, structure containing chemical and physicochemical information, bioactivity data for the related biological endpoint, or outcomes from clinical trials, etc. |
Overview of suggested general criteria to generate quality datasets useful in de novo design.
4.1 Balance
As discussed previously, several current data sets in the public domain are unbalanced due to the infrequent practice of reporting inactive compounds and negative data in general. Historically, the negative and inactive data of preclinical compounds has been ignored by most journals that favor the publication of most active compounds and positive results (
4.2 Confidence of the activity data
An unwritten rule on AI and computational projects in general is "garbage in, garbage out". This perspective has direct implications in drug design (
4.3 Chemical and structural diversity
In general, a compound dataset with a large or broad applicability domain, as captured by the diversity of the contents, can give rise to predictive models with a large coverage. This is, molecules from diverse chemical structures could be conveniently interpolated in those models. As a comparison in an experimental setting, high-throughput screening of chemical diverse libraries increases the chances to find hit compounds for targets for which no hit compounds have been previously identified.
Due to the rapid expansion of the chemical universe, recently called the ‘Big Bang’ of the chemical universe (
4.4 Preparation or curation
A general curation protocol used on drug discovery datasets is to eliminate duplicate structures, canonize their SMILES representation, eliminate salts, and metals. However, according to the main goal of the de novo design model, additional steps to prepare a dataset could be taking into account, for example: 1) eliminating compounds with structural PAINS to reduce the rate of false-positive compounds prediction; 2) deleting compounds reported with side effects and/or ADMET deficiencies, to prioritize the generation of safe and optimization compounds.; or 3) making sure to keep in the dataset compounds with high activity confidence to improve the quality of predicted outputs. This list must be adapted according to the main goal of the de novo design model. It is also noted the need to develop robust and consistent protocols that take into scout metal-containing compounds as they have a major role in medicinal inorganic chemistry (
4.5 Completeness
Chemical structures should contain the required or relevant information for the goals of the study. For instance, compounds should be annotated with stereochemistry information if the 3D structure and conformation is critical; electronic density and quantum chemical data if the reactivity is key point to predict; the type of the biological activity data such as biochemical, cell-based or functional assays; drug-drug interaction data, pharmacogenomics, or post-marketing annotations; should be aligned with the type of outcome to be predicted and later validated experimentally.
5 Perspectives of de novo design
One of the major perspectives of the de novo design is using balanced data sets (as much as experimental data is available) to build reliable models. Similar to QSAR predictive models, it is also crucial the validation of de novo protocols using standard and well-curated benchmark datasets (discussed in Section 3.6). With the increasing data availability to generate and train new models, it is becoming increasingly easy to explore regions of chemical space previously uncharted and continue contributing to the so-called “big bang” expansion of the chemical space. A major perspective in this direction is to explore biologically relevant compounds but outside the traditional small molecule chemical space (
6 Conclusion
Among the main types of datasets used in the novo design are on-demand collections, compounds annotated with biological activity, commercially available libraries, and natural products. More recently, a large benchmark data set was developed for machine learning applications. Although there is a general agreement in machine learning that the more data, the better, it is becoming more and more evident to consider the reliability and the quality of the data sets as critical features of the data. Part of the quality is associated with the balance between inactive and active compounds (in a rough analogy with the Yin-Yang concept), tasks that are not always feasible due to the general scarcity of negative (inactive compounds). The later point further emphasizes the continued need to publish and disclose negative results. Due to the fact that the experimental data of inactive compounds are not common, the community is using decoy data sets that by themselves are subject to design and refining using rational approaches. Decoy data sets try to fill the void of experimentally determined inactive molecules. Major criteria to take into account to generate compound data sets with high quality include balanced data sets in terms of active and inactive compounds (when the experimental information is available), structural and chemical diversity, curation or preparation according to the goals of the project, and complete information. All these together contribute to the perspectives of de novo design that foresees a continued and rapid expansion of molecules with the potential to become drugs.
Statements
Author contributions
All authors listed have made a substantial, direct, and intellectual contribution to the work and approved it for publication.
Funding
Authors are grateful to DGAPA, UNAM, Programa de Apoyo a Proyectos de Investigación e Innovación Tecnológica (PAPIIT), grant no. IN201321. We also thank the Dirección General de Cómputo y de Tecnologías de Información y Comunicación (DGTIC), UNAM, for the computational resources to use Miztli supercomputer at UNAM under project LANCAD-UNAM-DGTIC-335; and the innovation space UNAM-HUAWEI the computational resources to use their supercomputer under project-7 “Desarrollo y aplicación de algoritmos de inteligencia artificial para el diseño de fármacos aplicables al tratamiento de diabetes mellitus y cáncer”.
Acknowledgments
AC-H and EL-L are thankful to CONACyT, Mexico, for the Ph.D. scholarships number 847870 and 894234, respectively.
Conflict of interest
The author JLM-F declared that he was an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.
The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Abbreviations
2D/3D, two-dimensional/three-dimensional; ADMET, absorption, distribution, metabolism, excretion, and toxicity; AI, artificial intelligence; CADD, computer-aided drug design; COCONUT, Collection of Open NatUral ProdUcTs; DNN, deep neural networks; HBA, hydrogen bond acceptors; HBD, hydrogen bond donors; IMPPAT, A curated database of Indian Medicinal Plants, Phytochemistry And Therapeutics; MW, molecular weight; LBDD, ligand-based drug design; log P, octanol-water partition coefficient; NCE, new chemical entities; NIH(US), National Institutes of Health; PAINS, pan-assay interference compounds; Peru NPDB, Peruvian Natural Products Database; QSAR, quantitative structure-activity relationships; REAL, Enamine’s REadily AccessibLe; RNNs, recurrent neural networks; SBDD, structure-based drug design; TCM, Traditional Chinese Medicine; TPSA, topological surface area; UNPD, Universal Natural Product Database.
References
1
Altae-TranH.RamsundarB.PappuA. S.PandeV. (2017). Low data drug discovery with one-shot learning. ACS central Sci.3 (4), 283–293. 10.1021/acscentsci.6b00367
2
Arús-PousJ.PatronovA.BjerrumE. J.TyrchanC.ReymondJ. L.ChenH.et al (2020). SMILES-based deep generative scaffold decorator for de-novo drug design. J. cheminformatics12 (1), 38. 10.1186/s13321-020-00441-8
3
AtanasovA. G.ZotchevS. B.DirschV. M.SupuranC. T.International Natural Product Sciences Taskforce (2021). Natural products in drug discovery: Advances and opportunities. Nat. Rev. Drug Discov.20 (3), 200–216. 10.1038/s41573-020-00114-z
4
AwaleM.ReymondJ-L. (2015). Similarity mapplet: Interactive visualization of the directory of useful decoys and ChEMBL in high dimensional chemical spaces. J. Chem. Inf. Model.55 (8), 1509–1516. 10.1021/acs.jcim.5b00182
5
BajorathJ.Chávez-HernándezA. L.Duran-FrigolaM.Fernández-de GortariE.GasteigerJ.López-LópezE.et al (2022). Chemoinformatics and artificial intelligence colloquium: Progress and challenges in developing bioactive compounds. J. cheminformatics14 (1), 82. 10.1186/s13321-022-00661-0
6
BaliusT. E.AllenW. J.MukherjeeS.RizzoR. C. (2013). Grid-based molecular footprint comparison method for docking and de novo design: Application to HIVgp41. J. Comput. Chem.34 (14), 1226–1240. 10.1002/jcc.23245
7
Barazorda-CcahuanaH. L.RanillaL. G.Candia-PumaM. A.Cárcamo-RodriguezE. G.Centeno-LopezA. E.Davila-Del-CarpioG.et al (2023). PeruNPDB: The Peruvian natural products database for in silico drug screening. Sci. Rep.13 (1), 7577. 10.1038/s41598-023-34729-0
8
BauerM. R.IbrahimT. M.VogelS. M.BoecklerF. M. (2013). Evaluation and optimization of virtual screening workflows with DEKOIS 2.0-a public library of challenging docking benchmark sets. J. Chem. Inf. Model.53 (6), 1447–1462. 10.1021/ci400115b
9
BeatoC.BeccariA. R.CavazzoniC.LorenziS.CostantinoG. (2013). Use of experimental design to optimize docking performance: The case of LiGenDock, the docking module of LiGen, a new de novo design program. J. Chem. Inf. Model.53 (6), 1503–1517. 10.1021/ci400079k
10
BenderA.SchneiderN.SeglerM.Patrick WaltersW.EngkvistO.RodriguesT. (2022). Evaluation guidelines for machine learning tools in the chemical sciences. Nat. Rev. Chem.6 (6), 428–442. 10.1038/s41570-022-00391-9
11
BilodeauC.JinW.JaakkolaT.BarzilayR.JensenK. F. (2022). Generative models for molecular discovery: Recent advances and challenges. Comput. Mol. Sci.12 (5), e1608. 10.1002/wcms.1608
12
BlaschkeT.Arús-PousJ.ChenH.MargreitterC.TyrchanC.EngkvistO.et al (2020). Reinvent 2.0: An AI tool for de novo drug design. J. Chem. Inf. Model.60 (12), 5918–5922. 10.1021/acs.jcim.0c00915
13
BrownN.FiscatoM.SeglerM. H. S.VaucherA. C. (2019). GuacaMol: Benchmarking models for de Novo molecular design. J. Chem. Inf. Model.59 (3), 1096–1108. 10.1021/acs.jcim.8b00839
14
CaoL.GoreshnikI.CoventryB.CaseJ. B.MillerL.KozodoyL.et al (2020). De novo design of picomolar SARS-CoV-2 miniprotein inhibitors. Science370 (6515), 426–431. 10.1126/science.abd9909
15
Cereto-MassaguéA.GuaschL.VallsC.MuleroM.PujadasG.Garcia-VallvéS. (2012). DecoyFinder: An easy-to-use python GUI application for building target-specific decoy sets. Bioinformatics28 (12), 1661–1662. 10.1093/bioinformatics/bts249
16
Chávez-HernándezA. L.Medina-FrancoJ. L. (2023). Natural products subsets: Generation and characterization. Artif. Intell. Life Sci.3, 100066. 10.1016/j.ailsci.2023.100066
17
Chávez-HernándezA. L.Sánchez-CruzN.Medina-FrancoJ. L. (2020a). A fragment library of natural products and its comparative chemoinformatic characterization. Mol. Inf.39 (11), e2000050. 10.1002/minf.202000050
18
Chávez-HernándezA. L.Sánchez-CruzN.Medina-FrancoJ. L. (2020b). Fragment library of natural products and compound databases for drug discovery. Biomolecules10 (11), 1518. 10.3390/biom10111518
19
Chemriya (2023). CHEMriya. Available at: https://chemriya.com/ (accessed May 13, 2023).
20
Chemspace (2023). Freedom space. Available at: https://chem-space.com/compounds/freedom-space (accessed May 13, 2023).
21
ChenC. Y-C. (2011). TCM Database@Taiwan: The world’s largest traditional Chinese medicine database for drug screening in silico. PloS one6 (1), e15939. 10.1371/journal.pone.0015939
22
ChenX.LinY.LiuM.GilsonM. K. (2002). The binding database: Data management and interface design. Bioinformatics18 (1), 130–139. 10.1093/bioinformatics/18.1.130
23
CherkasovA. (2023). The ‘Big Bang’ of the chemical universe. Nat. Chem. Biol.19, 667–668. 10.1038/s41589-022-01233-x
24
CorsoG.StärkH.JingB.et al (2022). DiffDock: Diffusion steps, twists, and turns for molecular docking. arXiv [q-bio.BM]. Available at: http://arxiv.org/abs/2210.01776.
25
CostaR. P. O.LucenaL. F.SilvaL. M. A.ZocoloG. J.Herrera-AcevedoC.ScottiL.et al (2021). The SistematX web portal of natural products: An update. J. Chem. Inf. Model.61 (6), 2516–2522. 10.1021/acs.jcim.1c00083
26
DaviesM.NowotkaM.PapadatosG.DedmanN.GaultonA.AtkinsonF.et al (2015). ChEMBL web services: Streamlining access to drug discovery data and utilities. Nucleic acids Res.43 (W1), W612–W620. 10.1093/nar/gkv352
27
Dos Santos NascimentoI. J.de AquinoT. M.da Silva-JúniorE. F. (2021). Drug repurposing: A strategy for discovering inhibitors against emerging viral infections. Curr. Med. Chem.28 (15), 2887–2942. 10.2174/0929867327666200812215852
28
DRUGBANK (2023). Celecoxib. Available at: https://go.drugbank.com/drugs/DB00482 (accessed May 13, 2023).
29
Enamine (2023). Real database. Available at: https://enamine.net/compound-collections/real-compounds/real-database (accessed May 13, 2023).
30
EvansB. E.RittleK. E.BockM. G.DiPardoR. M.FreidingerR. M.WhitterW. L.et al (1988). Methods for drug discovery: Development of potent, selective, orally effective cholecystokinin antagonists. J. Med. Chem.31 (12), 2235–2246. 10.1021/jm00120a002
31
FourchesD.MuratovE.TropshaA. (2016). Trust, but verify II: A practical guide to chemogenomics data curation. J. Chem. Inf. Model.56 (7), 1243–1252. 10.1021/acs.jcim.6b00129
32
GalloK.KemmlerE.GoedeA.BeckerF.DunkelM.PreissnerR.et al (2023). SuperNatural 3.0-a database of natural products and natural product-based derivatives. Nucleic acids Res.51 (D1), D654–D659. 10.1093/nar/gkac1008
33
Gómez-BombarelliR.WeiJ. N.DuvenaudD.Hernández-LobatoJ. M.Sánchez-LengelingB.SheberlaD.et al (2018). Automatic chemical design using a data-driven continuous representation of molecules. ACS central Sci.4 (2), 268–276. 10.1021/acscentsci.7b00572
34
Gómez-GarcíaA.Medina-FrancoJ. L. (2022). Progress and impact of Latin American natural product databases. Biomolecules12 (9), 1202. 10.3390/biom12091202
35
GrebnerC. (2022). Webinar: "exploration and mining of large virtual chemical spaces. Available at: https://youtu.be/fMrI11SXwpU (accessed May 13, 2023).
36
GreenerJ. G.KandathilS. M.MoffatL.JonesD. T. (2022). A guide to machine learning for biologists. Nat. Rev. Mol. Cell Biol.23 (1), 40–55. 10.1038/s41580-021-00407-0
37
GrigalunasM.BrakmannS.WaldmannH. (2022). Chemical evolution of natural product structure. J. Am. Chem. Soc.144 (8), 3314–3329. 10.1021/jacs.1c11270
38
GuJ.GuiY.ChenL.YuanG.LuH. Z.XuX. (2013). Use of natural products as chemical library for drug discovery and network pharmacology. PloS one8 (4), e62839. 10.1371/journal.pone.0062839
39
Guo JJ.JanetJ. P.BauerM. R.NittingerE.GiblinK. A.PapadopoulosK.et al (2021). DockStream: A docking wrapper to enhance de novo molecular design. J. cheminformatics13 (1), 89. 10.1186/s13321-021-00563-7
40
Guo MM.ThostV.LiB.et al (2021). “Data-efficient graph grammar learning for molecular generation,” in International conference on learning representations, 9. February 2021. Available at: https://research.ibm.com/publications/data-efficient-graph-grammar-learning-for-molecular-generation (accessed May 13, 2023).
41
GuoM.ThostV.LiB.et al (2022). Data-efficient graph grammar learning for molecular generation. arXiv [cs.LG]. Available at: http://arxiv.org/abs/2203.08031.
42
HayesA.HunterJ. (2012). Why is publication of negative clinical trial data important?Br. J. Pharmacol.167 (7), 1395–1397. 10.1111/j.1476-5381.2012.02215.x
43
HuQ.PengZ.SuttonS. C.NaJ.KostrowickiJ.YangB.et al (2012). Pfizer global virtual library (PGVL): A chemistry design tool powered by experimentally validated parallel synthesis information. ACS Comb. Sci.14 (11), 579–589. 10.1021/co300096q
44
IBM (2022). How to use AI to discover new drugs and materials with limited data. Available at: https://research.ibm.com/blog/ai-discovery-with-limited-data#fnref-1 (accessed April 16, 2023).
45
IrwinJ. J. (2008). Community benchmarks for virtual screening. J. computer-aided Mol. Des.22 (3-4), 193–199. 10.1007/s10822-008-9189-4
46
JainA. N.NichollsA. (2008). Recommendations for evaluation of computational methods. J. computer-aided Mol. Des.22 (3-4), 133–139. 10.1007/s10822-008-9196-5
47
JumperJ.EvansR.PritzelA.GreenT.FigurnovM.RonnebergerO.et al (2021). Highly accurate protein structure prediction with AlphaFold. Nature596 (7873), 583–589. 10.1038/s41586-021-03819-2
48
JuskalianR.RegaladoA.OrcuttM.et al (2023). 10 breakthrough technologies 2020. Available at: https://www.technologyreview.com/10-breakthrough-technologies/2020/ (accessed February 26, 2020).
49
KadurinA.AliperA.KazennovA.MamoshinaP.VanhaelenQ.KhrabrovK.et al (2017). The cornucopia of meaningful leads: Applying deep adversarial autoencoders for new molecule development in oncology. Oncotarget8 (7), 10883–10890. 10.18632/oncotarget.14073
50
KimS.ChenJ.ChengT.GindulyteA.et al (2023). PubChem 2023 update. Nucleic acids Res.51 (D1), D1373–D1380. 10.1093/nar/gkac956
51
KoesD. R.CamachoC. J. (2012). ZINCPharmer: Pharmacophore search of the ZINC database. Nucleic acids Res.40, W409–W414. Web Server issue). 10.1093/nar/gks378
52
KorkmazS. (2020). Deep learning-based imbalanced data classification for drug discovery. J. Chem. Inf. Model.60 (9), 4180–4190. 10.1021/acs.jcim.9b01162
53
KornM.EhrtC.RuggiuF.GastreichM.RareyM. (2023). Navigating large chemical spaces in early-phase drug discovery. Curr. Opin. Struct. Biol.80, 102578. 10.1016/j.sbi.2023.102578
54
KramerC.LewisR. (2012). QSARs, data and error in the modern age of drug discovery. Curr. Top. Med. Chem.12 (17), 1896–1902. 10.2174/156802612804547380
55
KrennM.HäseF.NigamA.FriederichP.Aspuru-GuzikA. (2020). Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation. Mach. Learn. Sci. Technol.1 (4), 045024. 10.1088/2632-2153/aba947
56
KrishnanS. R.BungN.BulusuG.RoyA. (2021). Accelerating de novo drug design against novel proteins using deep learning. J. Chem. Inf. Model.61 (2), 621–630. 10.1021/acs.jcim.0c01060
57
KumarS. A.Ananda KumarT. D.BeerakaN. M.PujarG. V.SinghM.Narayana AkshathaH. S.et al (2022). Machine learning and deep learning in data-driven decision making of drug discovery and challenges in high-quality data acquisition in the pharmaceutical industry. Future Med. Chem.14 (4), 245–270. 10.4155/fmc-2021-0243
58
LeachA. R.GilletV. J. (2007). “Selecting diverse dets of compounds,” in An introduction to chemoinformatics (Dordrecht: Springer Netherlands), 119–139. 10.1007/978-1-4020-6291-9_6
59
LiS.WangL.MengJ.ZhaoQ.ZhangL.LiuH. (2022). De Novo design of potential inhibitors against SARS-CoV-2 Mpro. Comput. Biol. Med.147, 105728. 10.1016/j.compbiomed.2022.105728
60
LiY.ZhangL.LiuZ. (2018). Multi-objective de novo drug design with conditional graph generative model. J. cheminformatics10 (1), 33. 10.1186/s13321-018-0287-6
61
LiangY.FangR.RaoQ. (2022). An insight into the medicinal chemistry perspective of macrocyclic derivatives with antitumor activity: A systematic review. Molecules27 (9), 2837. 10.3390/molecules27092837
62
LipinskiC. A.LombardoF.DominyB. W.FeeneyP. J. (2001). Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings. Adv. drug Deliv. Rev.46 (1-3), 3–26. 10.1016/s0169-409x(00)00129-0
63
LiuX.YeK.van VlijmenH. W. T.IjzermanA. P.van WestenG. J. P. (2019). An exploration strategy improves the diversity of de novo ligands using deep reinforcement learning: A case for the adenosine A2A receptor. J. cheminformatics11 (1), 35. 10.1186/s13321-019-0355-6
64
López-LópezE.BajorathJ.Medina-FrancoJ. L. (2021a). Informatics for chemistry, biology, and biomedical sciences. J. Chem. Inf. Model.61 (1), 26–35. 10.1021/acs.jcim.0c01301
65
López-LópezE.Cerda-García-RojasC. M.Medina-FrancoJ. L. (2021b). Tubulin inhibitors: A chemoinformatic analysis using cell-based data. Molecules26 (9), 2483. 10.3390/molecules26092483
66
López-LópezE.Fernández-de GortariE.Medina-FrancoJ. L. (2022). Yes SIR! On the structure-inactivity relationships in drug discovery. Drug Discov. today27 (8), 2353–2362. 10.1016/j.drudis.2022.05.005
67
López-LópezE.Medina-FrancoJ. L. (2023). Towards decoding hepatotoxicity of approved drugs through navigation of multiverse and consensus chemical spaces. Biomolecules13 (1), 176. 10.3390/biom13010176
68
MaB.TerayamaK.MatsumotoS.IsakaY.SasakuraY.IwataH.et al (2021). Structure-based de novo molecular generator combined with artificial intelligence and docking simulations. J. Chem. Inf. Model.61 (7), 3304–3313. 10.1021/acs.jcim.1c00679
69
MaziarkaL.PochaA.KaczmarczykJ.RatajK.DanelT.WarchołM. (2020). Mol-CycleGAN: A generative model for molecular optimization. J. cheminformatics12 (1), 2. 10.1186/s13321-019-0404-1
70
Medina-FrancoJ. L.Flores-PadillaE. A.Chávez-HernándezA. L. (2022b). “Chapter 23 - discovery and development of lead compounds from natural sources using computational approaches,” in Evidence-based validation of herbal medicine. Editor MukherjeeP. K.Second Edition (Elsevier), 539–560. 10.1016/B978-0-323-85542-6.00009-3
71
Medina-FrancoJ. L.López-LópezE.AndradeE.Ruiz-AzuaraL.FreiA.GuanD.et al (2022a). Bridging informatics and medicinal inorganic chemistry: Toward a database of metallodrugs and metallodrug candidates. Drug Discov. today27 (5), 1420–1430. 10.1016/j.drudis.2022.02.021
72
Medina-FrancoJ. L.López-LópezE. (2022). The essence and transcendence of scientific publishing. Front. Res. metrics Anal.7, 822453. 10.3389/frma.2022.822453
73
Medina-FrancoJ. L.Martinez-MayorgaK.MeuriceN. (2014). Balancing novelty with confined chemical space in modern drug discovery. Expert Opin. drug Discov.9 (2), 151–165. 10.1517/17460441.2014.872624
74
Medina-FrancoJ. L.NavejaJ. J.López-LópezE. (2019). Reaching for the bright StARs in chemical space. Drug Discov. today24 (11), 2162–2169. 10.1016/j.drudis.2019.09.013
75
MendezD.GaultonA.BentoA. P.ChambersJ.De VeijM.FélixE.et al (2019). ChEMBL: Towards direct deposition of bioassay data. Nucleic acids Res.47 (D1), D930–D940. 10.1093/nar/gky1075
76
MohanrajK.KarthikeyanB. S.Vivek-AnanthR. P.ChandR. P. B.AparnaS. R.MangalapandiP.et al (2018). Imppat: A curated database of indian medicinal plants, phytochemistry and therapeutics. Sci. Rep.8 (1), 4329. 10.1038/s41598-018-22631-z
77
MouchlisV. D.AfantitisA.SerraA.FratelloM.PapadiamantisA. G.AidinisV.et al (2021). Advances in de novo drug design: From conventional to machine learning methods. Int. J. Mol. Sci.22 (4), 1676. 10.3390/ijms22041676
78
MysingerM. M.CarchiaM.IrwinJ. J.ShoichetB. K. (2012). Directory of useful decoys, enhanced (DUD-E): Better ligands and decoys for better benchmarking. J. Med. Chem.55 (14), 6582–6594. 10.1021/jm300687e
79
NewmanD. J.CraggG. M. (2020). Natural products as sources of new drugs over the nearly four decades from 01/1981 to 09/2019. J. Nat. Prod.83 (3), 770–803. 10.1021/acs.jnatprod.9b01285
80
NIH (2023). Natural product libraries. Available at: https://www.nccih.nih.gov/grants/natural-product-libraries.
81
NiitsuA.SugitaY. (2023). Towards de novo design of transmembrane α-helical assemblies using structural modelling and molecular dynamics simulation. Phys. Chem. Chem. Phys. PCCP25 (5), 3595–3606. 10.1039/d2cp03972a
82
NorinderU.NavejaJ. J.López-LópezE.MucsD.Medina-FrancoJ. L. (2019). Conformal prediction of HDAC inhibitors. SAR QSAR Environ. Res.30 (4), 265–277. 10.1080/1062936X.2019.1591503
83
Ntie-KangF.ZofouD.BabiakaS. B.MeudomR.ScharfeM.LifongoL. L.et al (2013). AfroDb: A select highly potent and diverse natural product library from african medicinal plants. PloS one8 (10), e78085. 10.1371/journal.pone.0078085
84
OlivecronaM.BlaschkeT.EngkvistO.ChenH. (2017). Molecular de-novo design through deep reinforcement learning. J. cheminformatics9 (1), 48. 10.1186/s13321-017-0235-x
85
OlmedoA. D.Medina-FrancoJ. L. (2020). “Chemoinformatic approach: The case of natural products of Panama,” in Cheminformatics and its applications (IntechOpen). 10.5772/intechopen.87779
86
OlmedoD. A.González-MedinaM.GuptaM. P.Medina-FrancoJ. L. (2017). Cheminformatic characterization of natural products from Panama. Mol. Divers.21 (4), 779–789. 10.1007/s11030-017-9781-4
87
PalazzesiF.PozzanA. (2022). “Deep learning applied to ligand-based de novo drug DesignDe novo drug design,” in Artificial intelligence in drug design. Editor HeifetzA. (New York, NY: Springer US), 273–299. 10.1007/978-1-0716-1787-8_12
88
PapadopoulosK.GiblinK. A.JanetJ. P.PatronovA.EngkvistO. (2021). De novo design with deep generative models based on 3D similarity scoring. Bioorg. Med. Chem.44, 116308. 10.1016/j.bmc.2021.116308
89
PatelH. M.NoolviM. N.SharmaP.JaiswalV.BansalS.LohanS.et al (2014). Quantitative structure–activity relationship (QSAR) studies as strategic approach in drug discovery. Med. Chem. Res. Int. J. rapid Commun. Des. Mech. action Biol. Act. agents23 (12), 4991–5007. 10.1007/s00044-014-1072-3
90
PerronQ.da SilvaV. B. R.AtwoodB.Gaston-MathéY. (2022b). Key points to succeed in Artificial Intelligence drug discovery projects. Chem. Int.44 (1), 19–21. 10.1515/ci-2022-0106
91
PerronQ.MirguetO.TajmouatiH.SkiredjA.RojasA.GohierA.et al (2022a). Deep generative models for ligand-based de novo design applied to multi-parametric optimization. J. Comput. Chem.43 (10), 692–703. 10.1002/jcc.26826
92
PilonA. C.ValliM.DamettoA. C.PintoM. E. F.FreireR. T.Castro-GamboaI.et al (2017). NuBBEDB: An updated database to uncover chemical and biological information from Brazilian biodiversity. Sci. Rep.7 (1), 7215. 10.1038/s41598-017-07451-x
93
Pilón-JiménezB. A.Saldívar-GonzálezF. I.Díaz-EufracioB. I.Medina-FrancoJ. L. (2019). Biofacquim: A Mexican compound database of natural products. Biomolecules9 (1), 31. 10.3390/biom9010031
94
PolykovskiyD.ZhebrakA.Sanchez-LengelingB.GolovanovS.TatanovO.BelyaevS.et al (2020). Molecular sets (MOSES): A benchmarking platform for molecular generation models. Front. Pharmacol.11, 565644. 10.3389/fphar.2020.565644
95
RéauM.LangenfeldF.ZaguryJ-F.LagardeN.MontesM. (2018). Decoys selection in benchmarking datasets: Overview and perspectives. Front. Pharmacol.9, 11. 10.3389/fphar.2018.00011
96
ReymondJ-L. (2015). The chemical space project. Accounts Chem. Res.48 (3), 722–730. 10.1021/ar500432k
97
RohrerS. G.BaumannK. (2009). Maximum unbiased validation (MUV) data sets for virtual screening based on PubChem bioactivity data. J. Chem. Inf. Model.49 (2), 169–184. 10.1021/ci8002649
98
SabeV. T.NtombelaT.JhambaL. A.MaguireG. E. M.GovenderT.NaickerT.et al (2021). Current trends in computer aided drug design and a highlight of drugs discovered via computational techniques: A review. Eur. J. Med. Chem.224, 113705. 10.1016/j.ejmech.2021.113705
99
Saldívar-GonzálezF. I.Aldas-BulosV. D.Medina-FrancoJ. L.PlissonF. (2022). Natural product drug discovery in the artificial intelligence era. Chem. Sci.13 (6), 1526–1546. 10.1039/d1sc04471k
100
Saldívar-GonzálezF. I.Medina-FrancoJ. L. (2022). Approaches for enhancing the analysis of chemical space for drug discovery. Expert Opin. drug Discov.17 (7), 789–798. 10.1080/17460441.2022.2084608
101
Saldívar-GonzálezF. I.ValliM.AndricopuloA. D.da Silva BolzaniV.Medina-FrancoJ. L. (2019). Chemical space and diversity of the NuBBE database: A chemoinformatic characterization. J. Chem. Inf. Model.59 (1), 74–85. 10.1021/acs.jcim.8b00619
102
Sánchez-CruzN.Pilón-JiménezB. A.Medina-FrancoJ. L. (2019) Functional group and diversity analysis of BIOFACQUIM: A Mexican natural product database. F1000Research8, Chem Inf Sci-2071. 10.12688/f1000research.21540.2
103
ScannellJ. W.BosleyJ.HickmanJ. A.DawsonG. R.TruebelH.FerreiraG. S.et al (2022). Predictive validity in drug discovery: What it is, why it matters and how to improve it. Nat. Rev. Drug Discov.21 (12), 915–931. 10.1038/s41573-022-00552-x
104
SchneiderG.ClarkD. E. (2019). Automated de novo drug design: Are we nearly there yet?Angew. Chem.58 (32), 10792–10803. 10.1002/anie.201814681
105
SchneiderG.Clément-ChomienneO.HilfigerL.SchneiderKirschBöhmet al (2000). Virtual screening for bioactive molecules by evolutionary de novo design. Angew. Chem.39 (22), 4130–4133. 10.1002/1521-3773(20001117)39:22<4130:aid-anie4130>3.0.co;2-e
106
SchneiderP.SchneiderG. (2017). Privileged structures revisited. Angew. Chem.56 (27), 7971–7974. 10.1002/anie.201702816
107
SchneiderP.WaltersW. P.PlowrightA. T.SierokaN.ListgartenJ.GoodnowR. A.et al (2020). Rethinking drug design in the artificial intelligence era. Nat. Rev. Drug Discov.19 (5), 353–364. 10.1038/s41573-019-0050-3
108
ScottiM. T.Herrera-AcevedoC.OliveiraT. B.CostaR. P. O.SantosS. Y. K. d. O.RodriguesR. P.et al (2018). SistematX, an online web-based cheminformatics tool for data management of secondary metabolites. Molecules23 (1), 103. 10.3390/molecules23010103
109
SheridanR. P. (2013). Time-split cross-validation as a method for estimating the goodness of prospective prediction. J. Chem. Inf. Model.53 (4), 783–790. 10.1021/ci400084k
110
ShipmanJ. T.SuX.HuaD.DesaireH. (2019). DecoyDeveloper: An on-demand, de novo decoy glycopeptide generator. J. proteome Res.18 (7), 2896–2902. 10.1021/acs.jproteome.9b00203
111
SimonovskyM.KomodakisN. (2018). “GraphVAE: Towards generation of small graphs using variational autoencoders,” in Artificial neural networks and machine learning – icann 2018 (Springer International Publishing), 2018, 412–422. 10.1007/978-3-030-01418-6_41
112
SkalicM.JiménezJ.SabbadinD.De FabritiisG. (2019b). Shape-based generative modeling for de novo drug design. J. Chem. Inf. Model.59 (3), 1205–1214. 10.1021/acs.jcim.8b00706
113
SkalicM.SabbadinD.SattarovB.SciabolaS.De FabritiisG. (2019a). From target to drug: Generative modeling for the multimodal structure-based ligand design. Mol. Pharm.16 (10), 4282–4291. 10.1021/acs.molpharmaceut.9b00634
114
SoaresT. A.Nunes-AlvesA.MazzolariA.RuggiuF.WeiG. W.MerzK. (2022). The (Re)-evolution of quantitative structure-activity relationship (qsar) studies propelled by the surge of machine learning methods. J. Chem. Inf. Model.62 (22), 5317–5320. 10.1021/acs.jcim.2c01422
115
SorokinaM.MerseburgerP.RajanK.YirikM. A.SteinbeckC. (2021). COCONUT online: Collection of open natural products database. J. cheminformatics13 (1), 2. 10.1186/s13321-020-00478-9
116
TingleB. I.TangK. G.CastanonM.GutierrezJ. J.KhurelbaatarM.DandarchuluunC.et al (2023). ZINC-22─A free multi-billion-scale database of tangible compounds for ligand discovery. J. Chem. Inf. Model.63 (4), 1166–1176. 10.1021/acs.jcim.2c01253
117
TongX.LiuX.TanX.JiangJ.XiongZ.et al (2021). Generative models for de novo drug design. J. Med. Chem.64 (19), 14011–14027. 10.1021/acs.jmedchem.1c00927
118
UllanatV. (2020). “Variational autoencoder as a generative tool to produce de-novo lead compounds for biological targets,” in 2020 14th international conference on innovations in information Technology (IIT), 102–107. 10.1109/IIT50501.2020.9299078
119
UNIIQUIM (2015). Uniiquim. Available at: https://uniiquim.iquimica.unam.mx/ (accessed May 13, 2023).
120
ValliM.dos SantosR. N.FigueiraL. D.NakajimaC. H.Castro-GamboaI.AndricopuloA. D.et al (2013). Development of a natural products database from the biodiversity of Brazil. J. Nat. Prod.76 (3), 439–444. 10.1021/np3006875
121
VamathevanJ.ClarkD.CzodrowskiP.DunhamI.FerranE.LeeG.et al (2019). Applications of machine learning in drug discovery and development. Nat. Rev. Drug Discov.18 (6), 463–477. 10.1038/s41573-019-0024-5
122
VeberD. F.JohnsonS. R.ChengH-Y.SmithB. R.WardK. W.KoppleK. D. (2002). Molecular properties that influence the oral bioavailability of drug candidates. J. Med. Chem.45 (12), 2615–2623. 10.1021/jm020017n
123
WangL.PangX.LiY.ZhangZ.TanW. (2017). Rader: A RApid DEcoy retriever to facilitate decoy based assessment of virtual screening. Bioinformatics33 (8), 1235–1237. 10.1093/bioinformatics/btw783
124
WangM.HsiehC-Y.WangJ.WangD.WengG.ShenC.et al (2022). Relation: A deep generative model for structure-based de novo drug design. J. Med. Chem.65 (13), 9478–9492. 10.1021/acs.jmedchem.2c00732
125
WarrW. A.NicklausM. C.NicolaouC. A.RareyM. (2022). Exploration of ultralarge compound collections for drug discovery. J. Chem. Inf. Model.62 (9), 2021–2034. 10.1021/acs.jcim.2c00224
126
WarrW. (2021). Report on an NIH workshop on ultralarge chemistry databases. Chemrxiv: 43. Available at: https://chemrxiv.org/engage/chemrxiv/article-details/60c75883bdbb89984ea3ada5.
127
WeiningerD. (1988). SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci.28 (1), 31–36. 10.1021/ci00057a005
128
WishartD. S.FeunangY. D.GuoA. C.LoE. J.MarcuA.GrantJ. R.et al (2018). DrugBank 5.0: A major update to the DrugBank database for 2018. Nucleic acids Res.46 (D1), D1074–D1082. 10.1093/nar/gkx1037
129
WishartD. S.KnoxC.GuoA. C.ChengD.ShrivastavaS.TzurD.et al (2008). DrugBank: A knowledgebase for drugs, drug actions and drug targets. Nucleic acids Res.36, D901–D906. 10.1093/nar/gkm958
130
WishartD. S.KnoxC.GuoA. C.ShrivastavaS.HassanaliM.StothardP.et al (2006). DrugBank: A comprehensive resource for in silico drug discovery and exploration. Nucleic acids Res.34, D668–D672. 10.1093/nar/gkj067
131
WuA.YeQ.ZhuangX.ChenQ.ZhangJ.WuJ.et al (2023a). Elucidating structures of complex organic compounds using a machine learning model based on the 13C NMR chemical shifts. Precis. Chem.1 (1), 57–68. 10.1021/prechem.3c00005
132
WuJ.XiaoY.CaiH.ZhaoD.LiY.et al (2023b). DeepCancerMap: A versatile deep learning platform for target- and cell-based anticancer drug discovery. Eur. J. Med. Chem.255, 115401. 10.1016/j.ejmech.2023.115401
133
WuZ.RamsundarB.FeinbergE. N.GomesJ.GeniesseC.PappuA. S.et al (2018). MoleculeNet: A benchmark for molecular machine learning. Chem. Sci.9 (2), 513–530. 10.1039/c7sc02664a
134
XieW.WangF.LiY.LaiL.PeiJ. (2022). Advances and challenges in de novo drug design using three-dimensional deep generative models. J. Chem. Inf. Model.62 (10), 2269–2279. 10.1021/acs.jcim.2c00042
135
YangJ.WangD.JiaC.WangM.HaoG.YangG. (2019). Freely accessible chemical database resources of compounds for in silico drug discovery. Curr. Med. Chem.26 (42), 7581–7597. 10.2174/0929867325666180508100436
136
YangX.YangG.ChuJ. (2023). The balanced matrix factorization for computational drug repositioning. arXiv [cs.CE]. Available at: http://arxiv.org/abs/2301.06448.
137
YuH. (2021). Responsible use of negative research outcomes-accelerating the discovery and development of new antibiotics. J. antibiotics74 (9), 543–546. 10.1038/s41429-021-00439-w
138
ZhangY.LuoM.WuP.WuS.LeeT. Y.BaiC. (2022). Application of computational biology and artificial intelligence in drug design. Int. J. Mol. Sci.23 (21), 13568. 10.3390/ijms232113568
139
ZhavoronkovA.IvanenkovY. A.AliperA.VeselovM. S.AladinskiyV. A.AladinskayaA. V.et al (2019). Deep learning enables rapid identification of potent DDR1 kinase inhibitors. Nat. Biotechnol.37 (9), 1038–1040. 10.1038/s41587-019-0224-x
Summary
Keywords
big data, chemoinformatics, chemical libraries, data quality, de novo design, drug discovery, machine learning, negative results
Citation
Chávez-Hernández AL, López-López E and Medina-Franco JL (2023) Yin-yang in drug discovery: rethinking de novo design and development of predictive models. Front. Drug Discov. 3:1222655. doi: 10.3389/fddsv.2023.1222655
Received
15 May 2023
Accepted
12 June 2023
Published
21 June 2023
Volume
3 - 2023
Edited by
Ho Leung Ng, Atomwise Inc., United States
Reviewed by
Rodolpho C. Braga, InsilicAll, Brazil
Jeremy J. Yang, University of New Mexico, United States
Updates

Check for updates
Copyright
© 2023 Chávez-Hernández, López-López and Medina-Franco.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: José L. Medina-Franco, medinajl@unam.mx
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.