Abstract
In recent years, the development of high-throughput screening (HTS) technologies and their establishment in an industrialized environment have given scientists the possibility to test millions of molecules and profile them against a multitude of biological targets in a short period of time, generating data in a much faster pace and with a higher quality than before. Besides the structure activity data from traditional bioassays, more complex assays such as transcriptomics profiling or imaging have also been established as routine profiling experiments thanks to the advancement of Next Generation Sequencing or automated microscopy technologies. In industrial pharmaceutical research, these technologies are typically established in conjunction with automated platforms in order to enable efficient handling of screening collections of thousands to millions of compounds. To exploit the ever-growing amount of data that are generated by these approaches, computational techniques are constantly evolving. In this regard, artificial intelligence technologies such as deep learning and machine learning methods play a key role in cheminformatics and bio-image analytics fields to address activity prediction, scaffold hopping, de novo molecule design, reaction/retrosynthesis predictions, or high content screening analysis. Herein we summarize the current state of analyzing large-scale compound data in industrial pharmaceutical research and describe the impact it has had on the drug discovery process over the last two decades, with a specific focus on deep-learning technologies.
Introduction
Digital data, in all shapes and sizes, are growing exponentially. According to the National Security Agency of the United States, the Internet is processing around 1.8 billion GB of data per day (). In 2011, digital information has grown nine times in volume in just 5 years () and by 2020, its amount in the world is expected to reach 35 trillion GB (). The recent development of deep learning and other artificial intelligence methods is fuelled by the desire to seek greater insight among the ever-increasing amount of data in several key industries and powered by technological advancements as in, for example, computer vision, natural language processing, internet of things (IoT), or computer hardware.
Over the past decade, there has been a remarkable increase in the amount of available compound activity, biomedical (; ; Schamberger et al., 2011), and genomics data (; ; Wilson and Nicholls, 2015) thanks to the rapid development of high-throughput screening (HTS) and gene sequencing technologies. Typically, databases in pharma companies contain around 1–4 million compounds with biological data for several thousands of biological end-points such as targets or activities in cellular assays. Furthermore, due to the increasing level of automation and standardization, larger data sets of consistent conditions have become available. All chemical compounds synthesized and/or extracted from publications represent around 96 million compounds (). Even though only a small fraction of them have associated biological information (Wang et al., 2014; ), these chemogenomics data sets alone already represent a formidable task for predictive modelling work.
The usage of new automation technologies resulted in a large volume of data, which has promoted the usage of machine learning (ML) methods. ML methods such as support vector machine (SVM), random forest (RF), or neural networks (NNs) have been used for data modelling in cheminformatics and bioinformatics for a long time. Only recently, various deep learning methods have become more popular due to the availability of large-scale training sets and high-performance computer hardware. An important difference between deep learning and previous ML methods is the flexibility of NN architectures and input/output data structures in deep learning methods and the automated extraction of features from raw data representations. This flexibility allows to design models that fit to the characteristics of the prediction problem (Wu et al., 2018; Xiong et al., 2019; Yang et al., 2019). Some of the popular NN architectures include convolutional NNs, recurrent NNs, autoencoders, and fully connected deep NNs. These deep learning methods have been applied (Ramsundar et al., 2017; ) on aspects of compound activity prediction (; ; ), de novo molecular design (), protein–ligand interaction prediction (; ), predictive toxicity (), and reaction prediction (Segler and Waller, 2017b). In this review, we will provide an overview on various types of large-scale data sets that are available in pharmaceutical industry. Such data sets offer a wealth of information that are unavailable in the public domain and give rise to a broad range of applications. Furthermore, we will exemplify the applications of artificial intelligence, in particular deep-learning technologies, that are powered through these large data sets on various problems in drug discovery.
Large-Scale Compound Data in Pharmaceutical Industry
The past two decades have seen an acceleration of compound data generation in pharmaceutical industry driven by the technical advancement of HTS (; ), parallel chemical synthesis (), as well as the by the introduction of automation in sequencing and imaging. The various types of large-scale compound data in pharmaceutical research are illustrated in Figure 1. A small molecule database belongs to the core infrastructure of industrial pharma R&D in order to store the results of lead identification and optimization campaigns, which are used for, e.g., structure–activity–relationship (SAR) analyses. The typical size of a compound collection at major pharma companies ranges from 1 to 4 million compounds (Schamberger et al., 2011; ). Compound activity data (including Administration Distribution Metabolism Excretion Toxicology (ADMET) end points) are the major part of the “Compound Data Estate” in pharmaceutical industry. Most of the SAR data come from the HTS campaigns carried out during the drug discovery projects, which typically comprise crude readouts generated from in vitro assays at single compound concentration—so called single-shot-potency—in the primary screening stage, and more accurate concentration response data (IC50s, EC50s, etc.) derived from multiple compound concentration experiments. Pharmaceutical databases allow for in-depth studies that may not be achievable with public data. Indeed, structuration and curation of private databases are done with the inclusion of concepts such as screening campaigns or lead optimization programs, which make possible a faster and easier analysis of high-quality data. Occasionally, the overall number of SAR data points in pharmaceutical companies was disclosed in the past; some numbers reported in literature are listed in Table 1. Although this information is not up-to-date, it can still give a sense of the scale of experimental compound data in pharmaceutical industry.
Figure 1
Table 1
| Company | # of SAR point | Date | Reference |
|---|---|---|---|
| AstraZeneca | 150 million single-shot SAR points, 14 milliona CR SAR points | Up to 2008 | (Proffitt, 2008; ) |
| Boehringer Ingelheim | 260 million single-shot SAR points, 7 million CR SAR points | Up to 2011 | () |
| Pfizer | 0.6 million CR SAR points | Up to 2005 | () |
| Johnson & Johnson | 30 million SAR points | Up to 2006 | () |
Number of SAR data point in large pharmaceutical companies reported in literatures.
a) This number includes external sources, up to 2012.
Comparing with conventional HTS screening with a limited number of data readouts per compound, high-content screening (HCS) () using automated microscopy generates images with multi-parameter readouts that provide an information-rich characterization of cellular phenotypic responses to small molecules. It has become an important tool for compound profiling and has led to a substantial increase in the amount of compound profiling data. For example, 460,800 images were produced through a screen comprising 100 384-well plates imaged with three fluorescent channels at four independent sites per well (). Hundreds of parameters can be extracted from each cell in the image quantifying information of morphological, geometric, intensity, and texture-based features. Recently Janssen reported (Simm et al., 2018) an image dataset for 524,371 compounds originally used for the detection of glucocorticoid receptor (GCR) nuclear translocation. For each cell in the image, 842 features were extracted, corresponding to roughly 440 million data points. The usage of image-based compound profiling data will be discussed in a subsequent section.
High throughput mRNA expression profiling can be used to characterize the response of cell culture models to perturbations such as small molecules acting as pharmacologic modulators (; ). These compounds induce transcriptional effects that can be used as gene signatures to discover new connections among compounds, pathways, and diseases. With one of these technologies, known as L1000™ Expression Profiling (profiling for 978 gene expressions) (; ), thousands of compounds can be screened per day at lower costs than conventional microarray techniques (Subramanian et al., 2017). Merck reported the screening of a set of 3,699 compounds using the Genometry L1000 platform to unveil a new target for compounds (). Janssen announced (; ) that they will use Genometry’s L1000 platform to generate gene-expression profiles for 250,000 compounds from Janssen’s small-molecule screening library. It is expected that more pharmaceutical companies will adopt similar technologies and approaches to generate large-scale transcriptomics data for compound profiling.
With the continuous increase in the amount and heterogeneity of data that are generated and stored in large repositories, the question of how to ensure and sustain data integrity gained more and more attention. The generation and storage of large amounts of data require significant investments in IT infrastructure. These investments are justified not only by efficiency gains for ongoing projects through elimination of manual steps to compile and analyze project-relevant data that ultimately lead to decisions on whether or not to pursue a certain molecule or compound class, but also perhaps even more so by the prospect to discover knowledge across projects as described for example in recent publications by Novartis (Wassermann et al., 2015a) or Boehringer Ingelheim (BI) (). All this is only possible if the data context is provided alongside the data itself, and when there is a profound understanding of the data quality. One important aspect for consideration is the assay technology that is applied for compound testing. The direct interference of compounds with an assay technology is a source for systematic errors, which should be considered when analyzing the respective data sets. In a recent example at BI (), the screening deck was assayed against an ion channel target for neuroprotection by means of a fluorometric imaging plate reader (FLIPR) assay (Sullivan et al., 1999). The screen yielded a high hit rate, and using a systematic overlap analysis with results from previous FLIPR campaigns, a large number of compounds most likely to be false positives were excluded from labor-intensive follow-up activities. Other important aspects regarding data quality are, for instance, compound purity, autofluorescence, or physicochemical properties such as aggregation propensity (), which can have a significant influence on assay results and need therefore to be taken into account as decision-relevant context. This can be accomplished by computational surrogate parameters or auxiliary experiments such as high-throughput solubility determination via nephelometry ().
Typically, data repositories within pharmaceutical companies evolve over years, and the best practices as to which data to store in such systems do so as well. This leads to situations in which legacy data are hardly comparable with present results, thereby limiting the chances to add value from mining data, which were generated at significantly different points in time. Efforts to set up data governance structures and to employ modern technologies around meta data management and central nomenclatures aim to address this issue and are currently underway in many companies (Proffitt, 2008).
Biological Profiling Descriptors for Hit Expansion
Traditionally, cheminformatic approaches focused on the use of molecular descriptors that are related to structure in order to describe the biological activities of compounds. Among them, structural fingerprints have been intensively used in similarity search, clustering, as well as in building SAR models (Willett, 2011). This is largely based on the hypothesis that structurally similar molecules are likely to bind to the same group of protein and then—as a consequence—share similar biological profiles (; ; Willett, 2011). In the late 1980s, NCI pioneered the implementation of a biological fingerprint to access the similarity of compounds (). In contrast to structural fingerprints, biological activity data are utilized to describe a compound, neglecting structural features. Furthermore, with the recent advent of phenotypic screening, we observe an increasing awareness that the cellular effects of a compound can be described by its interaction with the proteome, without requiring the knowledge of the molecular structure.
Efforts have been devoted to transpose various types of biological responses into fingerprint format that could be used to access biological similarity of ligands (; ; ; Plouffe et al., 2008; ). Recently, researchers of Novartis reported the use of the huge amount of in-house HTS data for this purpose (Petrone et al., 2012). The aggregated data from 195 biochemical and cell-based assays for around 1.5 million of compounds have been employed to generate biological fingerprints, so called HTS-FP. They stressed the usefulness in mixing biochemical and cell-based data in detecting molecules that can produce similar phenotype without necessarily presenting the same mode of action (Petrone et al., 2012). They demonstrated the complementarity between the HTS-FP and a state-of-the-art molecular fingerprint [e.g., ECFP4 (Rogers and Hahn, 2010)] in similarity searches, especially in relation to the scaffold hopping potential of HTS-FP to identify structurally diverse hits. On the other hand, biological fingerprints were found to be more efficient in a study related to screening plate selection and hit expansion (Petrone et al., 2012). Additionally, it was observed that biological fingerprint-based clusters contain compounds that interact with targets that operate jointly in the cell. In further work, the combination of HTS-FP with structural fingerprints via the use of various machine-learning approaches has showed promising results in HTS hit expansion (Riniker et al., 2014). Other studies showed the usefulness of HTS-FP for iterative screening purpose (). HTS-FP has one major drawback though, which is that predictions cannot be made for compounds that have not been previously tested in any HTS assays. In addition, HTS predominantly produces much more inactive than active, which consequently leads to quite sparse HTS-FP. To tackle these issues, have developed a method where missing bioactivity data were compensated by considering structural data in a so-called combined fingerprint (CESFP) (Figure 2). They reported a significant improvement when using CESFP compared to the use of HTS-FP and Extended Circular Fingerprints (ECFP) alone in random-forest based activity prediction models. This indicates a clear synergistic effect between structural and biological fingerprints. HTS-FP have also been employed for multitask ML. In a recent study, it was observed that HTS-FP and ECFP based activity predictions, while comparable in performance, could return hits containing different chemotypes, suggesting that combining these approaches can be an efficient way to explore the bioactive chemical space (Sturm et al., 2019).
Figure 2
Leveraging the transcriptional data such as gene expression profile (gene signature) in a cell could be another way to construct a biological profile descriptor. The publicly funded CMap database (; ) initially contained profiles of 164 drugs and later expanded to 1,309 FDA-approved small molecules. These small molecules were tested in five human cell lines, generating over 7,000 gene expression profiles in the database (). Compound induced gene signature profiles have been used for finding diverse hits () and drug repositioning (; Sirota et al., 2011). Although generating this kind of compound related cell perturbation data is still quite expensive, several pharmaceutical companies, as mentioned earlier, are moving in the direction of generating such data in a large scale. It can be expected that transcriptomics-based biological descriptors will be explored for hit identification in the future. Other biological descriptors derived from multiplexed image data have been reported and successfully used for several tasks, which will be discussed in the subsequent imaging section.
Analysis of Image-Based Profiling Data With Machine Learning
In the drug discovery process, biological imaging and image analysis are widely used at various stages ranging from preclinical research to clinical trials. Imaging techniques enable the visualization of phenotype and behavior at multiple levels, including full body of humans or animals, organs, tissues, cells, subcellular compartments, and single molecules. A wide range of available imaging techniques can help to reveal the distribution of a drug in the body, organ, and cell as well as its mechanism of action. Such techniques rely on image datasets obtained through automated microscopy. An example of a large-scale image dataset is given by The Cell Image Library (), which contains 919,265 five-channel fields of view related to 30,616 compounds. The most common imaging techniques are automated microscopy using several fluorescent markers as well as label free microscopy such as brightfield and digital phase contrast. These imaging techniques and the downstream data analysis produce a large amount of data and associated extracted features. For several decades, automatic analysis methods () have been successfully applied to identify objects such as organs, tissue types, cells, and subcellular compartments. Effects of diseases and drugs could be quantified by applying statistics and ML methods on the features that were extracted from the images in post-processing efforts. However, recent developments in deep NNs and specifically convolutional NNs (CNNs) are revolutionizing the field and setting new gold standards for key tasks such as segmentation and classification (; ; ; ). These new methods not only achieve better results but also avoid the time-consuming manual work of designing features and searching analysis methods for specific tasks. To achieve this, relatively large annotated data sets and substantial computational resources as provided in modern GPU clusters are required for training.
Deep neural nets (typically CNNs) have now been successfully applied for most tasks occurring in automated cell and tissue microscopy image analysis, including denoising (Su et al., 2015), super resolution (; ; Rivenson et al., 2018; Wang et al., 2019), stain normalization (), hit identification (Simm et al., 2018), protein localization (), cell cycle phase classification (), mechanism of action classification (), focus quality check (Yang et al., 2018), segmentation both in 2D and 3D (often using some version of a U-net architecture (Ronneberger et al., 2015)), and modality estimation (). Many tasks fall in the area of classification, including tasks such as quality control (Yang et al., 2018), object detection (Ren et al., 2017; ), or outcome classification (). Classification can be performed either on the image level or on the object level. In the latter case, it is linked to a localization or detection task to identify objects in a given image. One common two-step approach used is to first select candidate regions and then classify them. Alternatively, the network output consists of a probability map, which is analyzed in a postprocessing step to identify the objects. A typical architecture for classification is shown in Figure 3.
Figure 3
Since large amounts of annotated data are often not available for a specific task, strategies such as transfer learning are often applied, e.g., for classification tasks (
As mentioned above, HCS where cells are exposed to different compounds followed by automated multichannel microscopy and subsequent automatic feature extraction is producing much richer data for screening than traditional HTS. More advanced analysis of cells exposed to chemical perturbations allows to identify related spatial and temporal information. Different biological descriptors derived from multiplexed image data have been reported (
Predicting Compound Activity Using Large Chemogenomics Models
One of the main purposes of chemogenomics (
A major topic that has been briefly addressed previously is the necessity of data standardization and curation prior to building a predictive model. Chemical structures can be represented by different types of notations (SMILES, InChI, etc.) (
Several models (Wang et al., 2013; Sushko et al., 2014;
Table 2
| Ref. | Performance traditional ML | Performance deep-learning |
|---|---|---|
| ( | RF: MCC = 0.89 | DNN: MCC = 0.91 |
| ( | RF: AUC = 0.78 | MT NN: AUC = 0.82 |
| ( | SVM: MCC = 0.50, BEDROC = 0.88 | DNN_MC: MCC = 0.57, BEDROC = 0.92 |
| RF: MCC = 0.56, BEDROC = 0.82 | ||
| ( | SVM: AUC = 0.71 | ST: AUC = 0.72 |
| MT: AUC = 0.75 | ||
| ( | RF: Pearson = 0.783 | GNN: Pearson = 0.822 |
| (Segler and Waller, 2017b) | LR: Acc = 0.86 (reaction prediction) | NN: Acc = 0.92 (reaction prediction) |
| LR: Acc = 0.64 (retrosynthesis) | NN: Acc = 0.78 (retrosynthesis) | |
| (Wu et al., 2018) (3) | SVM: AUC = 0.822 | GC: AUC = 0.829 |
| (Xiong et al., 2019) (4) | SVM: AUC = 0.792 | Attentive FP: AUC = 0.832 |
| (Yang et al., 2019) (5) | RF: AUC = 0.619 | FFN: AUC = 0.788 |
| ( | RF: R2 = 0.42 | DNN: R2 = 0.49 |
| (Ramsundar et al., 2017) (7) | RF: R2 = 0.428 | ST: R2 = 0.448 |
| MT: R2 = 0.468 |
Performances comparison of traditional ML and DL in Drug Discovery.
LR, ST, MT, GC, GNN, and FFN refer to Linear Regression, Single- and Multi-Task, Graph Convolution, Graph, and Feedforward Neural Network, respectively. (1) Averaged performance on validation sets over 7 datasets. (2) Averaged performance on test sets over 19 datasets. (3) Performance on a test subset of the Tox21 dataset. (4) Performance on the HIV dataset. (5) Performance on the Tox21 dataset. (6) Averaged performance over 15 datasets. (7) Model performance on a test set.
Although it is crucial to have a sufficient amount of training data to infer target predictions, having high-quality data is also necessary. Indeed, available activity data can be erroneous due to the problematic nature of the compounds (
Other criterion to consider in HTS the druglikeness of a compound, which is determined by the compound’s physicochemical (PC) and toxicological properties. Various quality control pipelines created to filter out compounds employ straightforward filtering rules (
Very recently, a new consortium of pharmaceutical, technology, and academic partners has launched the “MELLODDY” (Machine Learning Ledger Orchestration for Drug Discovery) project (
Modelling Chemical Reactions From Large-Scale Synthesis Data
It is of crucial importance in drug discovery to be able to predict the feasibility of chemical reactions (
Figure 4

Process of reaction prediction on an exemplary target molecule [lidocaine (Reilly, 2009)]. Machine-learning methods are applied to, first, predict the synthetic feasibility of the molecule and, second, predict the chemical context leading to the best yield possible for the reaction.
In the following, we will focus on recent examples of predicting how to synthesize molecules by mining large corpora of experimental synthesis data. For more general reviews, we refer to recent publications (Warr, 2014;
Data Driven De Novo Molecule Design Through Generative Models and Data Augmentation
Even though industrial compound-bioactivity datasets have millions of data points, many assay results for specific compound series (typical for the lead optimization stage of a drug discovery project) have much less SAR data. However, these datasets can still be augmented and be further exploited with deep learning approaches, such as QSAR and generative modelling. Data augmentation is the process of adding noise or artificial perturbation to the samples in the dataset before training the model in order to make the final models more robust to overfitting (
Similar approaches have also been used in areas relevant to pharmaceutical research such as predicting concentrations of chemical compounds from spectroscopy data (
Figure 5

Canonical (A) and randomized (B) SMILES representations of Aspirin. Numbers represent the atom numberings assigned by the canonicalization algorithm (A) or randomized (B). Green arrows indicate how the molecular graph is traversed. Both SMILES strings represent the same molecule but, as the atom numbering changes, the generated SMILES strings do too. Figure extracted with permission from
A great surge of interest in cheminformatics applications of deep learning has happened in recent years when NNs were used to generate molecules represented by SMILES strings (
Figure 6

Sampling process of a pre-trained recurrent neural network. The generation process starts with a GO token, and at each step, the model computes a probability distribution of all possible characters. Then, the next character is sampled from it and fed back to predict the next character. The internal memory in the long short-term memory (LSTM) cells enables the predictions to take previous characters into account when generating the next character.
Data augmentation techniques have also been applied in molecular generative models. For example, they have shown to improve the quality of the chemical space generated in VAEs (
Deep-learning-based generative model has been applied successfully for prospective design of new druglike molecules with desired activities (
Conclusion
Over the past years, large amounts of heterogeneous data characterizing the biological action of small molecules have been accumulated in pharmaceutical R&D, stored in both proprietary and publicly available data bases. The origin of these data ranges from biochemical or cellular assays to experiments that investigate the impact of compounds on transcriptomics signatures and assays with imaging readouts. These fast-growing data have fuelled the application of data-savvy ML methods, and in particular deep learning, in order to detect patterns that allow to derive hypotheses for compound-mediated effects on biological (model) systems or to generate predictive models that can be employed at various stages during identification and optimization of new drug candidates. Together with deep-learning-based approaches to sample the drug-like chemical space that—depending on the use case—can be applied with or without predictions of synthetic accessibility, a plethora of potential high-impact applications is emerging. It offers the opportunity to accelerate early drug discovery and to enable a much more comprehensive exploration of the chemical space and the biological effects of its members than traditional wet lab and virtual screening approaches.
Funding
LD and JA-P have received funding from the European Union’s Horizon 2020 research and innovation program under the Marie Sklodowska Curie grant agreement No 676434, “Big Data in Chemistry” (“BIGCHEM”, http://bigchem.eu). The article reflects only the authors view and neither the European Commission nor the Research Executive Agency (REA) are responsible for any use that may be made of the information it contains.
Statements
Author contributions
JMK, BB, and HC wrote the section Large-Scale Compound Data in Pharmaceutical Industry. TK wrote the section Biological Profiling Descriptors for Hit Expansion. JK wrote the section Analysis of Image-Based Profiling Data With Machine Learning. LD wrote the section Predicting Compound Activity Using Large Chemogenomics Models. OE wrote the section Modelling Chemical Reactions From Large-Scale Synthesis Data. JA-P and EB wrote the section Data Driven de Novo Molecule Design Through Generative Models and Data Augmentation. LD and HC co-supervised the manuscript.
Conflict of interest
Authors LD, JA-P, JK, OE, EB, TK and HC were employed by AstraZeneca. Authors JMK and BB were employed by Boehringer Ingelheim Pharma GmbH & Co. KG.
References
1
AgrafiotisD. K.AlexS.DaiH.DerkinderenA.FarnumM.GatesP.et al. (2007). Advanced Biological and Chemical Discovery (ABCD): centralizing discovery knowledge in an inherently decentralized world. J. Chem. Inf. Model.47, 1999–2014. doi: 10.1021/ci700267w
2
Arús-PousJ.BlaschkeT.UlanderS.ReymondJ. L.ChenH.EngkvistO. (2019a). Exploring the GDB-13 chemical space using deep generative models. J. Cheminform.11, 20. doi: 10.1186/s13321-019-0341-z
3
Arús-PousJ.JohanssonS.PtykhodkoO.BjerrumE. J.TyrchanC.ReymondJ.-L. (2019b). Randomized SMILES strings improve the quality of molecular generative models. ChemRxiv Prepr. Available at: https://chemrxiv.org/articles/Randomized_SMILES_Strings_Improve_the_Quality_of_Molecular_Generative_Models/8639942/1 [Accessed July 5, 2019]. doi: 10.26434/chemrxiv.8639942.v2
4
BaellJ. B.HollowayG. A. (2010). New substructure filters for removal of pan assay interference compounds (PAINS) from screening libraries and for their exclusion in bioassays. J. Med. Chem.53, 2719–2740. doi: 10.1021/jm901137j
5
BaellJ. B.NissinkJ. W. M. (2018). Seven year itch: pan-assay interference compounds (PAINS) in 2017 - utility and limitations. ACS Chem. Biol.13, 36–44. doi: 10.1021/acschembio.7b00903
6
BarrattM. D.BasketterD. A.RobertsD. W. (1994). Skin sensitization structure-activity relationships for phenyl benzoates. Toxicol. Vitr.8, 823–826. doi: 10.1016/0887-2333(94)90077-9
7
BeckB. (2012). BioProfile—Extract knowledge from corporate databases to assess cross-reactivities of compounds. Bioorg. Med. Chem.20, 5428–5435. doi: 10.1016/j.bmc.2012.04.023
8
BeckB.SeeligerD.KrieglJ. M. (2015). The impact of data integrity on decision making in early lead discovery. J. Comput. Aided Mol. Des.29, 911–921. doi: 10.1007/s10822-015-9871-2
9
BickleM. (2010). The beautiful cell: high-content screening in drug discovery. Anal. Bioanal. Chem.398, 219–226. doi: 10.1007/s00216-010-3788-3
10
BjerrumE. J. (2017). SMILES enumeration as data augmentation for neural network modeling of molecules. ArXiv.
11
BjerrumE. J.GlahderM.SkovT. (2017). Data augmentation of spectral data for convolutional neural network (CNN) based deep chemometrics1–10.
12
BjerrumE. J.SattarovB. (2018). Improving chemical autoencoder latent space and molecular de novo generation diversity with heteroencoders. Biomolecules8, 131. doi: 10.3390/biom8040131
13
BlumL. C.ReymondJ. L. (2009). 970 Million druglike small molecules for virtual screening in the chemical universe database GDB-13. J. Am. Chem. Soc131, 8732–8733. doi: 10.1021/ja902302h
14
BohacekR. S.McMartinC.GuidaW. C. (2010). ChemInform abstract: the art and practice of structure-based drug design: a molecular modeling perspective. ChemInform27, no–no. doi: 10.1002/chin.199617316
15
BormanS. (1999). Reducing time to drug discovery. Chem. Eng. News77, 33–48. doi: 10.1021/cen-v077n010.p033
16
BoscN.AtkinsonF.FelixE.GaultonA.HerseyA.LeachA. R. (2019). Large scale comparison of QSAR and conformal prediction methods and their applications in drug discovery. J. Cheminform.11, 4. doi: 10.1186/s13321-018-0325-4
17
BoutrosM.HeigwerF.LauferC. (2015). Microscopy-based high-content screening. Cell163, 1314–1325. doi: 10.1016/J.CELL.2015.11.007
18
BrayM. A.GustafsdottirS. M.RohbanM. H.SinghS.LjosaV.SokolnickiK. L.et al. (2017). A dataset of images and morphological profiles of 30 000 small-molecule treatments using the Cell Painting assay. Gigascience6, 1–5. doi: 10.1093/gigascience/giw014
19
BreimanL. (2001). Random forests. Mach. Learn.45, 5–32. doi: 10.1023/A:1010933404324
20
BrenkR.SchipaniA.JamesD.KrasowskiA.GilbertI. H.FrearsonJ.et al. (2008). Lessons learnt from assembling screening libraries for drug discovery for neglected diseases. ChemMedChem3, 435–444. doi: 10.1002/cmdc.200700139
21
BrownN.FiscatoM.SeglerM. H. S.VaucherA. C. (2019). GuacaMol: benchmarking models for de novo molecular design. doi: 10.1021/acs.jcim.8b00839
22
CaicedoJ. C.CooperS.HeigwerF.WarchalS.QiuP.MolnarC.et al. (2017). Data-analysis strategies for image-based cell profiling. Nat. Methods14, 849–863. doi: 10.1038/nmeth.4397
23
CaronP. R.MullicanM. D.MashalR. D.WilsonK. P.SuM. S.MurckoM. A. (2001). Chemogenomic approaches to drug discovery. Chem. Biol.5, 464–470. Available at: http://www.ncbi.nlm.nih.gov/Genbank/genbankstats.html. [Accessed May 27, 2019]. doi: 10.1016/S1367-5931(00)00229-5
24
CarpenterA. E.JonesT. R.LamprechtM. R.ClarkeC.KangI.FrimanO.et al. (2006). CellProfiler: image analysis software for identifying and quantifying cell phenotypes. Genome Biol.7, R100. doi: 10.1186/gb-2006-7-10-r100
25
ChenC. L.MahjoubfarA.TaiL.-C.BlabyI. K.HuangA.NiaziK. R.et al. (2016). Deep learning in label-free cell classification. Sci. Rep.6, 21471. doi: 10.1038/srep21471
26
ChenH.EngkvistO.WangY.OlivecronaM.BlaschkeT. (2018). The rise of deep learning in drug discovery. Drug Discovery Today23, 1241–1250. doi: 10.1016/j.drudis.2018.01.039
27
ChoK.Van MerriënboerB.GulcehreC.BahdanauD.BougaresF.SchwenkH.et al. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. in EMNLP 2014 - 2014 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference, 1724–1734 doi: 10.3115/v1/D14-1179
28
ChristC. D.ZentgrafM.KrieglJ. M. (2012). Mining electronic laboratory notebooks: analysis, retrosynthesis, and reaction based enumeration. J. Chem. Inf. Model.52, 1745–1756. doi: 10.1021/ci300116p
29
ChristiansenE. M.YangS. J.AndoD. M.JavaherianA.SkibinskiG.LipnickS.et al. (2018). In silico labeling: predicting fluorescent labels in unlabeled images. Cell173, 792–803.e19. doi: 10.1016/j.cell.2018.03.040
30
CireşanD. C.GiustiA.GambardellaL. M.SchmidhuberJ. (2013). Mitosis detection in breast cancer histology images with deep neural networks. Berlin, Heidelberg: Springer, 411–418. doi: 10.1007/978-3-642-40763-5_51
31
ColeyC. W.BarzilayR.JaakkolaT. S.GreenW. H.JensenK. F. (2017). Prediction of organic reaction outcomes using machine learning. ACS Cent. Sci.3, 434–443. doi: 10.1021/acscentsci.7b00064
32
ColeyC. W.GreenW. H.JensenK. F. (2018). Machine learning in computer-aided synthesis planning. Acc. Chem. Res.51, 1281–1289. doi: 10.1021/acs.accounts.8b00087
33
Connectivity Map Available at: https://www.broadinstitute.org/connectivity-map-cmap [Accessed October 24, 2019].
34
CoreyE. J.Todd WipkeW. (1969). Computer-assisted design of complex organic syntheses. Science (80-.) 166, 178–192. doi: 10.1126/science.166.3902.178
35
Cortés-CirianoI.BenderA. (2019). Reliable prediction errors for deep neural networks using test-time dropout. J. Chem. Inf. Model.59, 3330–3339. doi: 10.1021/acs.jcim.9b00297
36
CortesC.VapnikV. (1995). Support vector networks machine active learning with applications to text classification. Mach. Learn.20, 273–297. doi: 10.1007/BF00994018
37
CummingJ. G.DavisA. M.MuresanS.HaeberleinM.ChenH. (2013). Chemical predictive modelling to improve compound quality. Nat. Rev. Drug Discovery12, 948–962. doi: 10.1038/nrd4128
38
DahlG. E.JaitlyN.SalakhutdinovR. (2014). Multi-task neural networks for QSAR Predictions. ArXiv. Available at: http://arxiv.org/abs/1406.1231 [Accessed September 25, 2019].
39
DahlinJ. L.NissinkJ. W. M.StrasserJ. M.FrancisS.HigginsL.ZhouH.et al. (2015). PAINS in the assay: chemical mechanisms of assay interference and promiscuous enzymatic inhibition observed during a sulfhydryl-scavenging HTS. J. Med. Chem.58, 2091–2113. doi: 10.1021/jm5019093
40
DaviesM.NowotkaM.PapadatosG.DedmanN.GaultonA.AtkinsonF.et al. (2015). ChEMBL web services: streamlining access to drug discovery data and utilities. Web Serv. Issue Publ. Online43, W612–W620. doi: 10.1093/nar/gkv352
41
De WolfH.De BondtA.TurnerH.GöhlmannH. W. (2016). Transcriptional characterization of compounds: lessons learned from the public LINCS data. Assay Drug Dev. Technol.14, 252–260. doi: 10.1089/adt.2016.715
42
DixonS. L.VillarH. O. (2010). ChemInform abstract: bioactive diversity and screening library selection via Affinity fingerprinting. ChemInform30, no–no. doi: 10.1002/chin.199916265
43
DürrO.SickB. (2016). Single-cell phenotype classification using deep convolutional neural networks. J. Biomol. Screen.21, 998–1003. doi: 10.1177/1087057116631284
44
EltonD. C.BoukouvalasZ.FugeM. D.ChungP. W. (2019). Deep learning for molecular design—a review of the state of the art. Mol. Syst. Des. Eng.4, 828–849. doi: 10.1039/c9me00039a
45
EngkvistO.NorrbyP.-O.SelmiN.LamY.PengZ.ShererE. C.et al. (2018). Computational prediction of chemical reactions: current status and outlook. Drug Discovery Today23, 1203–1218. doi: 10.1016/J.DRUDIS.2018.02.014
46
EulenbergP.KöhlerN.BlasiT.FilbyA.CarpenterA. E.ReesP.et al. (2017). Reconstructing cell cycle and disease progression using deep learning. Nat. Commun.8, 463. doi: 10.1038/s41467-017-00623-3
47
FeinbergE. N.SurD.WuZ.HusicB. E.MaiH.LiY.et al. (2018). PotentialNet for molecular property prediction. ACS Cent. Sci.4, 1520–1530. doi: 10.1021/acscentsci.8b00507
48
FengY.MitchisonT. J.BenderA.YoungD. W.TallaricoJ. A. (2009). Multi-parameter phenotypic profiling: using cellular effects to characterize small-molecule compounds. Nat. Rev. Drug Discovery8, 567–578. doi: 10.1038/nrd2876
49
FilzenT. M.KutchukianP. S.HermesJ. D.LiJ.TudorM. (2017). Representing high throughput expression profiles via perturbation barcodes reveals compound targets. PloS Comput. Biol.13, e1005335. doi: 10.1371/journal.pcbi.1005335
50
FliggeT. A.SchulerA. (2006). Integration of a rapid automated solubility classification into early validation of hits obtained by high throughput screening. J. Pharm. Biomed. Anal.42, 449–454. doi: 10.1016/j.jpba.2006.05.004
51
FliriA. F.LogingW. T.ThadeioP. F.VolkmannR. A. (2005a). Biological spectra analysis: Linking biological activity profiles to molecular structure. Proc. Natl. Acad. Sci. U. S. A.102, 261–266. doi: 10.1073/pnas.0407790101
52
FliriA. F.LogingW. T.ThadeioP. F.VolkmannR. A. (2005b). Biospectra analysis: Model proteome characterizations for linking molecular structure and biological response. J. Med. Chem.48, 6918–6925. doi: 10.1021/jm050494g
53
GaoH.StrubleT. J.ColeyC. W.WangY.GreenW. H.JensenK. F. (2018). Using machine learning to predict suitable conditions for organic reactions. ACS Cent. Sci.4, 1465–1476. doi: 10.1021/acscentsci.8b00357
54
GaultonA.HerseyA.NowotkaM.Patrícia BentoA.ChambersJ.MendezD.et al. (2016). The ChEMBL database in 2017. Nucleic Acids Res.45, 945–954. doi: 10.1093/nar/gkw1074
55
GawehnE.HissJ. A.SchneiderG. (2016). Deep learning in drug discovery. Mol. Inform.35, 3–14. doi: 10.1002/minf.201501008
56
Genometry Available at: https://www.linkedin.com/company/genometry-inc/about/ [Accessed October 24, 2019].
57
GilsonM. K.LiuT.BaitalukM.NicolaG.HwangL.ChongJ. (2015). BindingDB in 2015: a public database for medicinal chemistry, computational chemistry and systems pharmacology. Nucleic Acids Res.44, 1045–1053. doi: 10.1093/nar/gkv1072
58
GohG. B.SiegelC.VishnuA.HodasN. O.BakerN. (2017). Chemception: a deep neural network with minimal chemistry knowledge matches the performance of expert-developed QSAR/QSPR Models.
59
Gómez-BombarelliR.WeiJ. N.DuvenaudD.Hernández-LobatoJ. M.Sánchez-LengelingB.SheberlaD.et al. (2018). Automatic chemical design using a data-driven continuous representation of molecules. ACS Cent. Sci.4, 268–276. doi: 10.1021/acscentsci.7b00572
60
Gostardb. Available at: www.gostardb.com/gostar/.
61
GrisoniF.NeuhausC. S.GabernetG.MüllerA. T.HissJ. A.SchneiderG. (2018). Designing anticancer peptides by constructive machine learning. ChemMedChem13, 1300–1302. doi: 10.1002/cmdc.201800204
62
GuimaraesG. L.Sanchez-LengelingB.OuteiralC.FariasP. L. C.Aspuru-GuzikA. (2017). Objective-reinforced generative adversarial networks (ORGAN) for sequence generation models. doi: arXiv:1705.10843v3
63
GuyerM. S.CollinsF. S. (1995). How is the Human Genome Project doing, and what have we learned so far? Proc. Natl. Acad. Sci. U. S. A.92, 10841–10848. doi: 10.1073/pnas.92.24.10841
64
HellerS. R.McNaughtA.PletnevI.SteinS.TchekhovskoiD. (2015). InChI, the IUPAC international chemical identifier. J. Cheminform.7, 23. doi: 10.1186/s13321-015-0068-4
65
HertzbergR. P.PopeA. J. (2000). High-throughput screening: new technology for the 21st century. Curr. Opin. Chem. Biol.4, 445–451. doi: 10.1016/S1367-5931(00)00110-1
66
HochreiterS.SchmidhuberJ. (1997). Long short-term memory. Neural Comput.9, 1735–1780. doi: 10.1162/neco.1997.9.8.1735
67
HofmarcherM.RumetshoferE.ClevertD.-A.HochreiterS.KlambauerG. (2019). Accurate Prediction of Biological Assays with High-Throughput Microscopy Images and Convolutional Networks. J. Chem. Inf. Model.59, 1163–1171. doi: 10.1021/acs.jcim.8b00670
68
How library-scale gene-expression profiling is changing drug discovery Available at: https://www.statnews.com/sponsor/2017/02/17/library-scale-gene-expression-profiling-changing-drug-discovery/ [Accessed October 24, 2019].
69
HsiehJ.-H.SedykhA.HuangR.XiaM.TiceR. R. (2015). A data analysis pipeline accounting for artifacts in Tox21 quantitative high-throughput screening assays. J. Biomol. Screen.20, 887–897. doi: 10.1177/1087057115581317
70
HughesT. B.DangN.MillerG. P.SwamidassS. J. (2016). Modeling reactivity to biological macromolecules with a deep multitask network. ACS Cent. Sci.2, 529–537. doi: 10.1021/acscentsci.6b00162
71
Human Genome Project Results Available at: https://www.genome.gov/human-genome-project/results [Accessed October 24, 2019].
72
HungJ.RavelD.LopesS. C. P.RangelG.NeryO. A.MalleretB.et al. (2018). Applying faster R-CNN for object detection on malaria images. Available at: http://arxiv.org/abs/1804.09548 [Accessed June 20, 2019].
73
InChI and InChIKeys for chemical structures Available at: https://www.inchi-trust.org/ [Accessed October 24, 2019].
74
IorioF.RittmanT.GeH.MendenM.Saez-RodriguezJ. (2013). Transcriptional data: a new gateway to drug repositioning? Drug Discovery Today18, 350–357. doi: 10.1016/j.drudis.2012.07.014
75
Ishimatsu-TsujiY.SomaT.KishimotoJ. (2010). Identification of novel hair-growth inducers by means of connectivity mapping. FASEB J.24, 1489–1496. doi: 10.1096/fj.09-145292
76
JadhavA.FerreiraR. S.KlumppC.MottB. T.AustinC. P.IngleseJ.et al. (2010). Quantitative analyses of aggregation, autofluorescence, and reactivity artifacts in a screen for inhibitors of a thiol protease. J. Med. Chem.53, 37–51. doi: 10.1021/jm901070c
77
JanowczykA.BasavanhallyA.MadabhushiA. (2017). Stain normalization using sparse autoEncoders (StaNoSA): application to digital pathology. Comput. Med. Imaging Graph.57, 50–61. doi: 10.1016/j.compmedimag.2016.05.003
78
JinW.BarzilayR.JaakkolaT. (2018). Junction tree variational autoencoder for molecular graph generation. Available at: http://arxiv.org/abs/1802.04364 [Accessed September 26, 2019].
79
KauvarL. M.HigginsD. L.VillarH. O.SportsmanJ. R.Engqvist-GoldsteinÅ.BukarR.et al. (1995). Predicting ligand binding to proteins by affinity fingerprinting. Chem. Biol.2, 107–118. doi: 10.1016/1074-5521(95)90283-X
80
KeiserM. J.RothB. L.ArmbrusterB. N.ErnsbergerP.IrwinJ. J.ShoichetB. K. (2007). Relating protein pharmacology by ligand chemistry. Nat. Biotechnol.25, 197–206. doi: 10.1038/nbt1284
81
KensertA.HarrisonP. J.SpjuthO. (2019). Transfer learning with deep convolutional neural networks for classifying cellular morphological changes. SLAS Discovery Adv. Life Sci. R&D24, 466–475. doi: 10.1177/2472555218818756
82
KimS. (2016). Getting the most out of PubChem for virtual screening. Expert Opin. Drug Discovery11, 843–855. doi: 10.1080/17460441.2016.1216967
83
KimS.ChenJ.ChengT.GindulyteA.HeJ.HeS.et al. (2019a). PubChem 2019 update: improved access to chemical data. Nucleic Acids Res.47, D1102–D1109. doi: 10.1093/nar/gky1033
84
KingmaD. P.WellingM. (2013). Auto-encoding variational bayes. Available at: http://arxiv.org/abs/1312.6114 [Accessed September 26, 2019].
85
KnoxC.LawV.JewisonT.LiuP.LyS.FrolkisA.et al. (2011). DrugBank 3.0: a comprehensive resource for “omics” research on drugs. Nucleic Acids Res.39, D1035–D1041. doi: 10.1093/nar/gkq1126
86
KogejT.BlombergN.GreasleyP. J.MundtS.VainioM. J.SchambergerJ.et al. (2013). Big pharma screening collections: more of the same or unique libraries? the AstraZeneca–Bayer Pharma AG case. Drug Discovery Today18, 1014–1024. doi: 10.1016/J.DRUDIS.2012.10.011
87
KoutsoukasA.MonaghanK. J.LiX.HuanJ. (2017). Deep-learning: investigating deep neural networks hyper-parameters and comparison of performance to shallow methods for modeling bioactivity data. J. Cheminform.9, 42. doi: 10.1186/s13321-017-0226-y
88
KrausO. Z.BaJ. L.FreyB. J. (2016). Classifying and segmenting microscopy images with deep multiple instance learning. Bioinformatics32, i52–i59. doi: 10.1093/bioinformatics/btw252
89
KrausO. Z.GrysB. T.BaJ.ChongY.FreyB. J.BooneC.et al. (2017). Automated analysis of high-content microscopy data with deep learning. Mol. Syst. Biol.13, 924. doi: 10.15252/msb.20177551
90
LambJ.CrawfordE. D.PeckD.ModellJ. W.BlatI. C.WrobelM. J.et al. (2006). The Connectivity Map: Using Gene-Expression Signatures to Connect Small Molecules, Genes, and Disease. Science (80-. ) 313, 1929–1935. doi: 10.1126/science.1132939
91
LaufkötterO.SturmN.BajorathJ.ChenH.EngkvistO. (2019). Combining structural and bioactivity-based fingerprints improves prediction performance and scaffold-hopping capability. chemRxiv.11, 54. doi: 10.26434/chemrxiv.7725209.v1
92
LenselinkE. B.Ten DijkeN.BongersB.PapadatosG.Van VlijmenH. W. T.KowalczykW.et al. (2017). Beyond the hype: deep neural networks outperform established methods using a ChEMBL bioactivity benchmark set. J. Cheminform.9, 45. doi: 10.1186/s13321-017-0232-0
93
LinA. I.MadzhidovT. I.KlimchukO.NugmanovR. I.AntipinI. S.VarnekA. (2016). Automatized assessment of protective group reactivity: a step toward big reaction data analysis. J. Chem. Inf. Model.56, 2140–2148. doi: 10.1021/acs.jcim.6b00319
94
LiuK.SunX.JiaL.MaJ.XingH.WuJ.et al. (2019). Chemi-net: a molecular graph convolutional network for accurate drug property prediction. Int. J. Mol. Sci.20, 3389. doi: 10.3390/ijms20143389
95
LooL.-H.WuL. F.AltschulerS. J. (2007). Image-based multivariate profiling of drug responses from single cells. Nat. Methods4, 445–453. doi: 10.1038/nmeth1032
96
MaJ.SheridanR. P.LiawA.DahlG. E.SvetnikV. (2015). Deep neural nets as a method for quantitative structure-activity relationships. J. Chem. Inf. Model.55, 263–274. doi: 10.1021/ci500747n
97
MacarronR.BanksM. N.BojanicD.BurnsD. J.CirovicD. A.GaryantesT.et al. (2011). Impact of high-throughput screening in biomedical research. Nat. Rev. Drug Discovery10, 188–195. doi: 10.1038/nrd3368
98
MartinE. J.PolyakovV. R.ZhuX.-W.TianL.MukherjeeP.LiuX. (2019). All-Assay-Max2 pQSAR: Activity Predictions as Accurate as Four-Concentration IC50s for 8558 Novartis Assays. J. Chem. Inf. Model. doi: 10.1021/acs.jcim.9b00375
99
MartinY. C.KofronJ. L.TraphagenL. M. (2002). Do structurally similar molecules have similar biological activity? J. Med. Chem.45, 4350–4358. Available at: http://www.ncbi.nlm.nih.gov/pubmed/12213076 [Accessed June 20, 2019]. doi: 10.1021/jm020155c
100
MayrA.KlambauerG.UnterthinerT.HochreiterS. (2016). DeepTox: toxicity prediction using deep learning. Front. Environ. Sci.3, 80. doi: 10.3389/fenvs.2015.00080
101
MayrA.KlambauerG.UnterthinerT.SteijaertM.WegnerJ. K.CeulemansH.et al. (2018). Large-scale comparison of machine learning methods for drug target prediction on ChEMBL. Chem. Sci.9, 5441–5451. doi: 10.1039/C8SC00148K
102
MayrL. M.BojanicD. (2009). Novel trends in high-throughput screening. Curr. Opin. Pharmacol.9, 580–588. doi: 10.1016/j.coph.2009.08.004
103
MELLODDY Consortium| Available at: https://cordis.europa.eu/project/rcn/223634/factsheet/en [Accessed October 24, 2019]
104
MerkD.FriedrichL.GrisoniF.SchneiderG. (2018). De novo design of bioactive small molecules by artificial intelligence. Mol. Inform.37, 1700153. doi: 10.1002/minf.201700153
105
MervinL. H.AfzalA. M.DrakakisG.LewisR.EngkvistO.BenderA. (2015). Target prediction utilising negative bioactivity data covering large chemical space. J. Cheminform.7, 51. doi: 10.1186/s13321-015-0098-y
106
MüllerA. T.HissJ. A.SchneiderG. (2018). Recurrent neural network model for constructive peptide design. J. Chem. Inf. Model.58, 472–479. doi: 10.1021/acs.jcim.7b00414
107
MuresanS.PetrovP.SouthanC.KjellbergM. J.KogejT.TyrchanC.et al. (2011). Making every SAR point count: the development of chemistry connect for the large-scale integration of structure and bioactivity data. Drug Discovery Today16, 1019–1030. doi: 10.1016/j.drudis.2011.10.005
108
NehmeE.WeissL. E.MichaeliT.ShechtmanY. (2018). Deep-STORM: super-resolution single-molecule microscopy by deep learning. Optica5, 458. doi: 10.1364/OPTICA.5.000458
109
OlivecronaM.BlaschkeT.EngkvistO.ChenH. (2017). Molecular de-novo design through deep reinforcement learning. J. Cheminform.9, 48. doi: 10.1186/s13321-017-0235-x
110
OuyangW.AristovA.LelekM.HaoX.ZimmerC. (2018). Deep learning massively accelerates super-resolution localization microscopy. Nat. Biotechnol.36, 460–468. doi: 10.1038/nbt.4106
111
PaoliniG. V.ShaplandR. H. B.van HoornW. P.MasonJ. S.HopkinsA. L. (2006). Global mapping of pharmacological space. Nat. Biotechnol.24, 805–815. doi: 10.1038/nbt1228
112
ParicharakS.IJzermanA. P.BenderA.NigschF. (2016). Analysis of iterative screening with stepwise compound selection based on novartis in-house HTS data. ACS Chem. Biol.11, 1255–1264. doi: 10.1021/acschembio.6b00029
113
PärnamaaT.PartsL. (2017). Accurate classification of protein subcellular localization from high-throughput microscopy images using deep learning. Genes|Genomes|Genetics7, 1385–1392. doi: 10.1534/g3.116.033654
114
PascaleC. (2015). Genometry Announces Deal with Janssen for Library-Scale Gene-Expression Profiling | Business Wire. Available at: https://www.businesswire.com/news/home/20151007006618/en#.VhZdNWTBzRZ [Accessed June 20, 2019].
115
PaulK. D.ShoemakerR. H.HodesL.MonksA.ScudieroD. A.RubinsteinL.et al. (1989). Display and analysis of patterns of differential activity of drugs against human tumor cell lines: development of mean graph and COMPARE algorithm. J. Natl. Cancer Inst.81, 1088–1092. doi: 10.1093/jnci/81.14.1088
116
PearceB. C.SofiaM. J.GoodA. C.DrexlerD. M.StockD. A. (2006). An empirical process for the design of high-throughput screening deck filters. J. Chem. Inf. Model.46, 1060–1068. doi: 10.1021/ci050504m
117
PetroneP. M.SimmsB.NigschF.LounkineE.KutchukianP.CornettA.et al. (2012). Rethinking molecular similarity: comparing compounds on the basis of biological activity. ACS Chem. Biol.7, 1399–1409. doi: 10.1021/cb3001028
118
Pharma Companies Join Forces to Train AI for Drug Discovery Collectively Available at: https://www.biopharmatrend.com/post/97-pharma-companies-join-forces-to-train-ai-for-drug-discovery-collectively/ [Accessed June 5, 2019].
119
PlouffeD.BrinkerA.McNamaraC.HensonK.KatoN.KuhenK.et al. (2008). In silico activity profiling reveals the mechanism of action of antimalarials discovered in a high-throughput screen. Proc. Natl. Acad. Sci.105, 9059–9064. doi: 10.1073/pnas.0802982105
120
PolykovskiyD.ZhebrakA.Sanchez-LengelingB.GolovanovS.TatanovO.BelyaevS.et al. (2018a). Molecular sets (MOSES): a benchmarking platform for molecular generation models.
121
PolykovskiyD.ZhebrakA.VetrovD.IvanenkovY.AladinskiyV.MamoshinaP.et al. (2018b). Entangled conditional adversarial autoencoder for de novo drug discovery. Mol. Pharm.15, 4398–4405. doi: 10.1021/acs.molpharmaceut.8b00839
122
ProffittA. (2008). AstraZeneca invests in data, discovery management - bio-IT World. Available at: http://www.bio-itworld.com/issues/2008/july-august/best-practices-astrazeneca.html [Accessed June 20, 2019].
123
PrykhodkoO.JohanssonS.KotsiasP.-C.BjerrumE. J.EngkvistO.ChenH. (2019). A de novo molecular generation method using latent vector based generative adversarial network. doi: 10.26434/chemrxiv.8299544.v1
124
PutinE.AsadulaevA.IvanenkovY.AladinskiyV.Sanchez-LengelingB.Aspuru-GuzikA.et al. (2018). Reinforced adversarial neural computer for de novo molecular design. J. Chem. Inf. Model.58, 1194–1204. doi: 10.1021/acs.jcim.7b00690
125
Pyzer-KnappE. O. (2018). Bayesian optimization for accelerated drug discovery. IBM J. Res. Dev.62, 2, 1–2:7. doi: 10.1147/JRD.2018.2881731
126
RamsundarB.LiuB.WuZ.VerrasA.TudorM.SheridanR. P.et al. (2017). Is multitask deep learning practical for pharma? J. Chem. Inf. Model.57, 2068–2076. doi: 10.1021/acs.jcim.7b00146
127
Reaxys Database. Available at: https://www.reaxys.com/#/login [Accessed October 24, 2019].
128
ReillyT. J. (2009). The preparation of lidocaine. J. Chem. Educ.76, 1557. doi: 10.1021/ed076p1557
129
ReisenF.Sauty de ChalonA.PfeiferM.ZhangX.GabrielD.SelzerP. (2015). Linking phenotypes and modes of action through high-content screen fingerprints. Assay Drug Dev. Technol.13, 415–427. doi: 10.1089/adt.2015.656
130
RenS.HeK.GirshickR.SunJ. (2017). Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell.39, 1137–1149. doi: 10.1109/TPAMI.2016.2577031
131
ReymondJ.-L. (2015). The chemical space project. Acc. Chem. Res.48, 722–730. doi: 10.1021/ar500432k
132
RinikerS.WangY.JenkinsJ. L.LandrumG. A. (2014). Using information from historical high-throughput screens to predict active compounds. J. Chem. Inf. Model.54, 1880–1891. doi: 10.1021/ci500190p
133
RivensonY.GöröcsZ.GünaydınH.ZhangY.WangH.OzcanA.et al., (2018). “Conference on lasers and electro-optics,” in deep learning microscopy: enhancing resolution, field-of-view and depth-of-field of optical microscopy images using neural networks (Washington, D.C: OSA), AM1J.5. doi: 10.1364/CLEO_AT.2018.AM1J.5
134
RogersD.HahnM. (2010). Extended-connectivity fingerprints. J. Chem. Inf. Model.50, 742–754. doi: 10.1021/ci100050t
135
RonnebergerO.FischerP.BroxT. (2015). U-Net: convolutional networks for biomedical image segmentation. Cham: Springer, 234–241. doi: 10.1007/978-3-319-24574-4_28
136
SchambergerJ.GrimmM.SteinmeyerA.HillischA. (2011). Rendezvous in chemical space? Comparing the small molecule compound libraries of bayer and schering. Drug Discovery Today16, 636–641. doi: 10.1016/j.drudis.2011.04.005
137
SchneiderG.FechnerU. (2005). Computer-based de novo design of drug-like molecules. Nat. Rev. Drug Discovery4, 649–663. doi: 10.1038/nrd1799
138
SchneiderN.LoweD. M.SayleR. A.TarselliM. A.LandrumG. A. (2016). Big data from pharmaceutical patents: a computational analysis of medicinal chemists’ bread and butter. J. Med. Chem.59, 4385–4402. doi: 10.1021/acs.jmedchem.6b00153
139
SchreckJ. S.ColeyC. W.BishopK. J. M. (2019). Learning Retrosynthetic Planning through Simulated Experience. ACS Cent. Sci.5, 970–981. doi: 10.1021/acscentsci.9b00055
140
SchwallerP.GaudinT.LányiD.BekasC.LainoT. (2018a). “Found in translation”: predicting outcomes of complex organic chemistry reactions using neural sequence-to-sequence models. Chem. Sci.9, 6091–6098. doi: 10.1039/c8sc02339e
141
SchwallerP.LainoT.GaudinT.BolgarP.BekasC.LeeA. A. (2018b). Molecular Transformer - a model for uncertainty-calibrated chemical reaction prediction. Available at: http://arxiv.org/abs/1811.02633 [Accessed June 25, 2019].
142
SciFinder. Available at: https://scifinder.cas.org [Accessed October 24, 2019]
143
SeglerM. H. S.KogejT.TyrchanC.WallerM. P. (2018a). Generating focused molecule libraries for drug discovery with recurrent neural networks. ACS Cent. Sci.4, 120–131. doi: 10.1021/acscentsci.7b00512
144
SeglerM. H. S.PreussM.WallerM. P. (2018b). Planning chemical syntheses with deep neural networks and symbolic AI. Nature555, 604–610. doi: 10.1038/nature25978
145
SeglerM. H. S.WallerM. P. (2017a). Modelling chemical reasoning to predict and invent reactions. Chem. A Eur. J.23, 6118–6128. doi: 10.1002/chem.201604556
146
SeglerM. H. S.WallerM. P. (2017b). Neural-symbolic machine learning for retrosynthesis and reaction prediction. Chem. A Eur. J.23, 5966–5971. doi: 10.1002/chem.201605499
147
SilvermanB. W.JonesM. C. (1989). E. Fix and J.L. Hodges (1951): An Important contribution to nonparametric discriminant analysis and density estimation: commentary on fix and hodges (1951). Int. Stat. Rev./Rev. Int. Stat.57, 233. doi: 10.2307/1403796
148
SimmJ.KlambauerG.AranyA.SteijaertM.WegnerJ. K.GustinE.et al. (2018). Repurposing high-throughput image assays enables biological activity prediction for drug discovery. Cell Chem. Biol.25, 611–618.e3. doi: 10.1016/j.chembiol.2018.01.015
149
SirotaM.DudleyJ. T.KimJ.ChiangA. P.MorganA. A.Sweet-CorderoA.et al. (2011). Discovery and preclinical validation of drug indications using compendia of public gene expression data. Sci. Transl. Med.3, 96ra77–96ra77. doi: 10.1126/scitranslmed.3001318
150
SterlingT.IrwinJ. J. (2015). ZINC 15 – Ligand discovery for everyone. J. Chem. Inf. Model.55, 2324–2337. doi: 10.1021/acs.jcim.5b00559
151
StorkC.ChenY.ŠíchoM.KirchmairJ. (2019). Hit Dexter 2.0: Machine-learning models for the prediction of frequent hitters. J. Chem. Inf. Model.59, 1030–1043. doi: 10.1021/acs.jcim.8b00677
152
StorkC.WagnerJ.FriedrichN. O.de Bruyn KopsC.ŠíchoM.KirchmairJ. (2018). Hit dexter: a machine-learning model for the prediction of frequent hitters. ChemMedChem13, 564–571. doi: 10.1002/cmdc.201700673
153
SturmN.SunJ.VandriesscheY.MayrA.KlambauerG.CarlssonL.et al. (2019). Application of bioactivity profile-based fingerprints for building machine learning models. J. Chem. Inf. Model.59, 962–972. doi: 10.1021/acs.jcim.8b00550
154
SuH.XingF.KongX.XieY.ZhangS.YangL. (2015). “Robust Cell Detection and Segmentation in Histopathological Images Using Sparse Reconstruction and Stacked Denoising Autoencoders,” in Medical image computing and computer-assisted intervention: MICCAI. International Conference on Medical Image Computing and Computer-Assisted Intervention. 383–390. doi: 10.1007/978-3-319-24574-4_46
155
SubramanianA.NarayanR.CorselloS. M.PeckD. D.NatoliT. E.LuX.et al. (2017). A next generation connectivity map: L1000 platform and the first 1,000,000 Profiles. Cell171, 1437–1452.e17. doi: 10.1016/j.cell.2017.10.049
156
SullivanE.TuckerE. M.DaleI. L. (1999). “Calcium signaling protocols,” in measurement of [Ca<sup<2+</sup>]; Using the fluorometric imaging plate reader (FLIPR) (New Jersey: Humana Press), 125–134. doi: 10.1385/1-59259-250-3:125
157
SunJ.JeliazkovaN.ChupakinV.Golib-DzibJ. F.EngkvistO.CarlssonL.et al. (2017). ExCAPE-DB: An integrated large scale dataset facilitating big data analysis in chemogenomics. J. Cheminform.9, 1–9. doi: 10.1186/s13321-017-0203-5
158
SushkoI.SalminaE.PotemkinV. A.PodaG.TetkoI. V. (2012). ToxAlerts: A web server of structural alerts for toxic chemicals and compounds with potential adverse reactions. J. Chem. Inf. Model.52, 2310–2316. doi: 10.1021/ci300245q
159
SushkoY.NovotarskyiS.KörnerR.VogtJ.AbdelazizA.TetkoI. V. (2014). Prediction-driven matched molecular pairs to interpret QSARs and aid the molecular optimization process. J. Cheminform.6, 1–18. doi: 10.1186/s13321-014-0048-0
160
TennantR. W.AshbyJ. (1991). Classification according to chemical structure, mutagenicity to Salmonella and level of carcinogenicity of a further 39 chemicals tested for carcinogenicity by the U.S. National Toxicology Program. Mutat. Res. Genet. Toxicol.257, 209–227. doi: 10.1016/0165-1110(91)90002-D
161
ThomsonReuters. Available at: https://www.thomsonreuters.com/en.html [Accessed October 24, 2019].
162
TsubakiM.TomiiK.SeseJ. (2019). Compound-protein interaction prediction with end-to-end learning of neural networks for graphs and sequences. Bioinformatics35, 309–318. doi: 10.1093/bioinformatics/bty535
163
WangH.RivensonY.JinY.WeiZ.GaoR.GünaydınH.et al. (2019). Deep learning enables cross-modality super-resolution in fluorescence microscopy. Nat. Methods16, 103–110. doi: 10.1038/s41592-018-0239-0
164
WangL.MaC.WipfP.LiuH.SuW.XieX.-Q. (2013). TargetHunter: an in silico target identification tool for predicting therapeutic potential of small organic molecules based on chemogenomic database. AAPS J.15, 395–406. doi: 10.1208/s12248-012-9449-z
165
WangY.SuzekT.ZhangJ.WangJ.HeS.ChengT.et al. (2014). PubChem BioAssay: 2014 update. Nucleic Acids Res.42, D1075–D1082. doi: 10.1093/nar/gkt978
166
WarrW. A. (2014). A short review of chemical reaction database systems, computer-aided synthesis design, reaction prediction and synthetic feasibility. Mol. Inform.33, 469–476. doi: 10.1002/minf.201400052
167
WassermannA. M.LounkineE.DaviesJ. W.GlickM.CamargoL. M. (2015a). The opportunities of mining historical and collective data in drug discovery. Drug Discovery Today20, 422–434. doi: 10.1016/j.drudis.2014.11.004
168
WassermannA. M.LounkineE.HoepfnerD.Le GoffG.KingF. J.StuderC.et al. (2015b). Dark chemical matter as a promising starting point for drug lead discovery. Nat. Chem. Biol.11, 958–966. doi: 10.1038/nchembio.1936
169
WeiningerD. (1988). SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Model.28, 31–36. doi: 10.1021/ci00057a005
170
WeiningerD.WeiningerA.WeiningerJ. L. (1989). SMILES. 2. Algorithm for generation of unique SMILES notation. J. Chem. Inf. Comput. Sci.29, 97–101. doi: 10.1021/ci00062a008
171
WillettP. (2011). Similarity-based data mining in files of two-dimensional chemical structures using fingerprint measures of molecular resemblance. Wiley Interdiscip. Rev. Data Min. Knowl. Discovery1, 241–251. doi: 10.1002/widm.26
172
WilsonB. J.NichollsS. G. (2015). The human genome project, and recent advances in personalized genomics. Risk Manage. Healthc. Policy8, 9–20. doi: 10.2147/RMHP.S58728
173
WuZ.RamsundarB.FeinbergE. N.GomesJ.GeniesseC.PappuA. S.et al. (2018). MoleculeNet: a benchmark for molecular machine learning. Chem. Sci.9, 513–530. doi: 10.1039/c7sc02664a
174
XiongZ.WangD.LiuX.ZhongF.WanX.LiX.et al. (2019). Pushing the boundaries of molecular representation for drug discovery with the graph attention mechanism. J. Med. Chem. acs.jmedchem.9b00959. doi: 10.1021/acs.jmedchem.9b00959
175
XuY.LinK.WangS.WangL.CaiC.SongC.et al. (2019). Deep learning for molecular generation. Future Med. Chem.11, 567–597. doi: 10.4155/fmc-2018-0358
176
YangJ. J.UrsuO.LipinskiC. A.SklarL. A.OpreaT. I.BologaC. G. (2016). Badapple: promiscuity patterns from noisy evidence. J. Cheminform.8, 29. doi: 10.1186/s13321-016-0137-3
177
YangK.SwansonK.JinW.ColeyC.EidenP.GaoH.et al. (2019). Analyzing learned molecular representations for property prediction. J. Chem. Inf. Model.59, 3370–3388. doi: 10.1021/acs.jcim.9b00237
178
YangS. J.BerndlM.Michael AndoD.BarchM.NarayanaswamyA.ChristiansenE.et al. (2018). Assessing microscope image focus quality with deep learning. BMC Bioinf.19, 77. doi: 10.1186/s12859-018-2087-4
179
YoshidaM.HinkleyT.TsudaS.Abul-HaijaY. M.McburneyR. T.KulikovV.et al. (2018). Exploring sequence space for antimicrobial peptides using evolutionary algorithms and machine learning. available at: https://blogit.itu.dk/evoblissproject/wp-content/uploads/sites/19/2018/03/yoshida_2018_preprint_Using-Evolutionary-Algorithms-and-Machine-Learning-to-Explore-Sequence-Space-for-the-Discovery-of-Antimicrobial-Peptides_.pdf [Accessed August 2, 2019].
180
YouJ.LiuB.YingR.PandeV.LeskovecJ. (2018). Graph convolutional policy network for goal-directed molecular graph generation. Available at: http://arxiv.org/abs/1806.02473 [Accessed September 26, 2019].
181
YoungD. W.BenderA.HoytJ.McWhinnieE.ChirnG.-W.TaoC. Y.et al. (2008). Integrating high-content screening and ligand-target prediction to identify mechanism of action. Nat. Chem. Biol.4, 59–68. doi: 10.1038/nchembio.2007.53
182
ZhaiY.ChenK.ZhongY.ZhouB.AinscowE.WuY.-T.et al. (2016). An automatic quality control pipeline for high-throughput screening hit identification. J. Biomol. Screen.21, 832–841. doi: 10.1177/1087057116654274
183
ZhangH. (2004). “Proceedings of the seventeenth international florida artificial intelligence research society conference, FLAIRS 2004,” in the optimality of Naive Bayes, 562–567. Available at: https://www.aaai.org/Papers/FLAIRS/2004/Flairs04-097.pdf [Accessed September 25, 2019].
184
ZhangW.LiR.ZengT.SunQ.KumarS.YeJ.et al. (2015). Deep model based transfer and multi-task learning for biological image analysis in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1475–1484 doi: 10.1145/2783258.2783304
185
ZhangY.LeeA. A. (2019). Bayesian semi-supervised learning for uncertainty-calibrated prediction of molecular properties and active learning. Chem. Sci.10, 8154–8163. doi: 10.1039/c9sc00616h
186
ZhavoronkovA.IvanenkovY. A.AliperA.VeselovM. S.AladinskiyV. A.AladinskayaA. V.et al. (2019). Deep learning enables rapid identification of potent DDR1 kinase inhibitors. Nat. Biotechnol.37, 1038–1040. doi: 10.1038/s41587-019-0224-x
Summary
Keywords
Artificial intelligence, deep learning, Chemogenomics, Large-scale data, pharmaceutical industry
Citation
David L, Arús-Pous J, Karlsson J, Engkvist O, Bjerrum EJ, Kogej T, Kriegl JM, Beck B and Chen H (2019) Applications of Deep-Learning in Exploiting Large-Scale and Heterogeneous Compound Data in Industrial Pharmaceutical Research. Front. Pharmacol. 10:1303. doi: 10.3389/fphar.2019.01303
Received
07 August 2019
Accepted
14 October 2019
Published
05 November 2019
Volume
10 - 2019
Edited by
Jianfeng Pei, Peking University, China
Reviewed by
Alexander Sedykh, Sciome LLC, United States; Maxim Kuznetsov, Insilico Medicine, Inc., United States
Updates

Check for updates
Copyright
© 2019 David, Arús-Pous, Karlsson, Engkvist, Bjerrum, Kogej, Kriegl, Beck and Chen.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Laurianne David, Laurianne.david1@gmail.com; Hongming Chen, Hongming.Chen71@hotmail.com
This article was submitted to Experimental Pharmacology and Drug Discovery, a section of the journal Frontiers in Pharmacology
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.