Abstract
We independently analyzed two large public domain datasets that contain 1H-NMR spectral data from lung cancer and sex studies. The biobanks were sourced from the Karlsruhe Metabolomics and Nutrition (KarMeN) study and Bayesian Automated Metabolite Analyzer for NMR data (BATMAN) study. Our approach of applying novel artificial intelligence (AI)-based algorithms to NMR is an attempt to globalize metabolomics and demonstrate its clinical applications. The intention of this study was to analyze the resulting spectra in the biobanks via AI application to demonstrate its clinical applications. This technique enables metabolite mapping in areas of localized enrichment as a measure of true activity while also allowing for the accurate categorization of phenotypes.
1. Introduction
The field of metabolomics is the most recent addition to the “-Omics” discipline. The core objective of this emerging field is to record all metabolites within a biological sample. Metabolites are understood to be by-products of cellular metabolism with a weight of ~2 kDa or less (, ). Water-soluble metabolites have the ability to communicate with the environment and the microbiome due to the mobility around the open biological system (). Consequently, metabolomics is essential for “systems biology” due to its particular scope analogous to fields such as genomics and proteomics (). “Hence, genomics and proteomics identify what could happen, metabolomics identifies what is currently happening in a system” (). The metabolomics framework is capable of examining endogenous metabolites and signal molecules that are by-products or participate in gene regulation, protein function, and enzymatic activity. Based on these, we identify ‘true activity' as a representation of what is currently happening in a biological system (). Additionally, metabolomics is often a consequence of “exposomics”, which is a series of factors that include diet, lifestyle, pollutants, medication, and the microbiome itself (Figure 1A) (). It is particularly valuable as it is capable of capturing the thousands of small molecule interactions within a given organism (). Therefore, a significant portion of research has been invested in the potential of tracking the downregulation and upregulation patterns of metabolites or biomarkers in order to interpret fluctuations in biological function (, ).
Figure 1
Broadly speaking, there are two metabolomics methodologies: The first is targeted metabolomics, which establishes associations between defined metabolites and known phenotypic states (
A potential workhorse instrumentation for untargeted metabolomics integration is nuclear magnetic resonance (NMR) due to its holistic detection capability combined with high sensitivity (though not as high as mass spectrometry) for low molecular weight biomarkers. It is typical to use NMR and mass spectrometry (MS) in tandem with multivariate analysis (
In the case of NMR, the standardized workflow generates thousands of signals which include true signals from metabolites, adducts, and fragments, as well as noise signals from contaminants and artifacts (
There is an abundance of applications that have demonstrated that AI is not a one size fits all; therefore, one must borrow and hybridize concepts from genome-wide association studies (GWASs) and Mummichog in an attempt to map all possible metabolite matches to a pathway via mass spectroscopy, solely focusing on regions of localized enrichment as they are assumed to be a reflection of “true activity” (
Our approach involves harnessing global metabolomics in addition to multivariate analysis in tandem with NMR to investigate metabolites and their correlation with sex and lung cancer. In this study, we use the data provided by two large biobank databases. All data relating to sex were curated and analyzed by Rist et al. (
2. Materials and methods
2.1. Data collection
For this investigation, we obtained open-source datasets from the health study by Rist et al. (
The KarMeN study (
Padayachee et al. (
The strict inclusion/exclusion parameters and the handling of samples in both studies gave us confidence in the integrity and excellence of both datasets, thus enabling us to perform our own analysis. The inputs we availed of were solely that of 1H-NMR datasets.
2.2. Data processing
In one-dimensional 1H-NMR spectroscopy, the signals are represented as the frequency domain resulting from the Fourier transform of a time-domain signal. These are given in units of parts per million (ppm), which is pre-determined at 0.0 ppm based on the chemical shift reference. Data processing was performed prior to any analysis to ensure the integrity and reliability of the results.
For the Padayachee et al. (
Regarding the Rist et al. (
Furthermore, the resulting pre-processing steps from the studies by Rist et al. (
The data obtained from the study by Padayachee et al. (
As per common practice in NMR, we removed water and its corresponding ppm as this often accounts for the majority of peak intensity and can mask minor variations in the NMR spectra. Due to the difference in obtained data, standardization was required, whereby the negative values within the dataset were set to zero and mean-centered scaling was applied to the Rist et al. (
Finally, the dataset was divided into two sets: a test set comprising 33% of the data and a training set with 66% of the data. This partitioning ensures an unbiased evaluation of the algorithm's performance. To determine the significance of different features in the dataset, the widely adopted statistical test known as the ANOVA F-test was employed for feature selection. In order to comprehensively evaluate the algorithm, a 10-fold cross-validation technique was applied. This method is commonly employed in machine learning to assess the algorithm's performance across multiple subsets of the dataset. By dividing the data into 10 equal parts, the algorithm was trained and evaluated 10 times, each time using a different combination of nine parts for training and one part for testing. This approach provides a more robust assessment of the algorithm's generalization capability and overall performance.
3. Results
The data were generated by obtaining open-source datasets from the Rist et al. (
We tested the integrity of our outputs by comparing them to the published analyses of the original datasets (
3.1. Lung cancer case study
Our analysis of the data provided from the Bayesian Automated Metabolite Analyzer lung cancer study (
Figure 2

Heatmap of leading features in (A) lung cancer cohorts and in (B) health and sexes. This heatmap is a representation of the top features and the correlations relative to other features. The feature was determined by a singular NMR unit (bin or bucket), measured in units of chemical shift (ppm). The location of the ppm was determined by ANOVA F-values. The features found through NMR analysis of plasma can be used to categorize the (A) lung cancer metabolome and (B) among sexes and determine the states of health.
Figure 3

Graphical outputs visualizing the linear relationship between ppm. (A) Minimum spanning tree (Mst) generated by the Fruchterman–Reingold algorithm used to visualize all ppm in the healthy category with correlations above a 90% threshold. Nodes closer together in the center have a stronger correlation and nodes far apart around the perimeter have little to no correlation. (B) Mst used to visualize all ppm in the diseased category with correlations above a 90% threshold.
Figure 4

(A) Boxplots demonstrating the significance of changes between healthy controls and lung cancer groups and between male–female cohorts. The boxplot demonstrates the absolute difference between the means of each feature. These features were further analyzed and identified to be the following metabolites; 2-aminoisobutyric acid, dimethylmalonic acid, tartaric acid, and glycine. These identified metabolites were among the lead features used to categorize the phenotypes of interest. Green represents the healthy controlled cohort, while red represents the lung cancer cohort. The binned NMR spectral data from the Padayachee et al. (
Figure 5

Kernel density plot used to visualize the distribution of lung cancer and the distribution of male–female cohorts. The above scatter plots demonstrate a clear separation among the cohorts. (A) For the lung cancer cohorts, the features of interest include dimethylmalonic, tartaric acid, glycine, and acetone. (B) For the distribution of sexes, the features of interest include creatinine 1, creatine 1, 2-hydroxy-2-methylbutyric (HMB) acid, and valine 1.
Figure 6

Graphical outputs visualizing the linear relationship between ppm. (A) Minimum spanning tree (Mst) generated by the Fruchterman–Reingold algorithm used to visualize all ppm in the male (blue) category with correlations above a 90% threshold. Nodes closer together in the center have a stronger correlation and nodes far apart around the perimeter have little to no correlation. (B) Mst used to visualize all ppm in the female (purple) category with correlations above a 90% threshold.
Figure 2A is a heatmap of leading features in lung cancer cohorts. The leading 20 metabolites contained in this heatmap are essential for characterizing phenotypic states. Of these 20, we have found asparagine, creatine, glycerol, threonine, glucose, citrate, and lactate. Moreover, we have identified tartaric acid, which was not on the list of key metabolites in the Padayachee et al. (
Our in silico analysis provided the following: Figures 3A and B are graphical outputs to visualize metabolomic relationships distilled down from a total of approximately 2 million relationships. The distillation of these relationships is further represented in Figures 4A and 5A which highlight the variability in the top-ranking metabolites. In summary, we have funneled down the key metabolites involved in lung cancer.
3.2. KarMeN health analysis among sexes
Our analysis of the data provided from the Karlsruhe Metabolomics and Nutrition study (
Figure 2B is a heatmap of leading features in the determination of sex in healthy cohorts. The leading 20 metabolites contained in this heatmap are essential for characterizing phenotypic states. Of these 20, we have found creatinine, creatine, glycerol, glycine, sarcosine, isoleucine, and valine. Moreover, we have identified 2-hydroxy-2-methylbutyric (HMB) acid, which was not in the list of key metabolites in the Rist et al. (
Figures 6A and B are graphical outputs to visualize metabolomic relationships distilled down from a total of approximately 2 million relationships. The distillation of these relationships is further represented in Figures 4B, 5B, which highlight the variability in the top-ranking metabolites. In summary, we have funneled down the key metabolites involved in distinguishing sex in healthy people.
4. Discussion
The primary objective of this study was to analyze the human metabolome in the plasma by way of globalized metabolomics profiling by harnessing 1H-NMR, to determine the factors that significantly impact the metabolic profile of a healthy cohort compared to a lung cancer cohort, and to distinguish the variables among the sexes. Therefore, we performed our study and established a strict in silico experimental standardization, which we applied to data structuring, data treatment, and post-analysis treatments. When collecting open-source data, we ensured that all sample collections were standardized in terms of fasting, collection time points, and general pre-analysis handling. We also searched for healthy datasets with strict exclusion and inclusion criteria that excluded groups that suffered from acute or chronic diseases or were on medication, as we wanted a dataset that represented “true health,” thereby decreasing variation. In contrast, the medication and acute/chronic disease exclusion criteria cannot be applied to the lung cancer cohort as they must undergo medical treatment in tandem with the study. Furthermore, this fundamental difference may be one variable that explains the variability when testing the integrity of the algorithm. Through additional analysis, we found that our process is capable of generating high-integrity categorization with minimal variation. The difference among predictive capabilities per dataset could be due to the number of samples; n = 301 (
Furthermore, some AI algorithms may require a relatively small amount of data to achieve satisfactory results, while others, particularly deep learning algorithms, often benefit from large-scale datasets. The size of the dataset required is directly proportional to the type of AI used and its field of application. Even a large dataset may not be useful if it is noisy, incomplete, or biased. A primary issue is the problem of complex, highly specialized, and specific fields focusing on molecular interactions, protein structures, or drug discovery that typically require domain expertise and specialized knowledge. As a result, the problem space is more constrained, and the available data may be more targeted and focused. In such cases, a smaller sample size can still provide meaningful insights and accurate predictions.
The impact of our analytical approach can be found in Figure 4. Many of our leading 20 metabolites have significant overlap with the pre-existing analysis (
We recognize that there are requirements for additional analysis and broadening of the inclusion criteria. Participants that are obese and/or smoking must be included and recorded for an accurate representation of the healthy population, as studies demonstrate that nicotine does have neuroprotective qualities (
Owing to the fact that NMR metabolomics provides a quantitative and holistic view of all of the metabolites contained, there is no reason that this technology cannot be applied to other diseases. In this article, we have successfully harnessed AI and metabolomic techniques to broaden the search parameters that aid in a comprehensive understanding of disease and wellbeing. The advancements made here can offer a snapshot of the entire biological system, which allows us to ascertain an accurate understanding of the phenotype in question, paving the way for true precision medicine.
5. Conclusion
From our analyses of NMR spectra from two separate biobanks, we have established that our approach has direct clinical applications. Our approach of harnessing AI and NMR to globalize metabolomics enables us to identify metabolites, to highlight them as regions of localized enrichment as a measure of true activity, while enabling us to accurately categorize phenotypes of interest.
Statements
Data availability statement
Publicly available datasets were analyzed in this study. This data can be found at: Padayachee et al. (
Ethics statement
Ethical review and approval was not required for the study on human participants in accordance with the local legislation and institutional requirements. The patients/participants provided their written informed consent to participate in this study.
Author contributions
LS and KHM initially discussed the potential of this research. LS, BM, and SB were involved in the coding and statistical evaluation of the data. LS, BM, and KHM wrote the manuscript. All authors contributed to the article and approved the submitted version.
Funding
The authors thank Enterprise Ireland (EI) for their mentorship and encouragement.
Acknowledgments
The authors thank Alsessor, an AI accelerator program sponsored by the Tangent of Trinity College Dublin, for initial mentoring and encouragement.
Conflict of interest
LS, BM, and SB was employed by the Meta-Flux Ltd.
The remaining author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fmed.2023.1162808/full#supplementary-material
References
1.
JohnsonCIvanisevicJSiuzdakG. Metabolomics: beyond biomarkers and towards mechanisms. Nat Rev Mol Cell Biol. (2016) 17:451–9. 10.1038/nrm.2016.25
2.
BegerRDDunnWSchmidtMAGrossSSKirwanJACascanteM. Metabolomics enables precision medicine: A White Paper, Community perspective. Metabolomics. (2016) 12:149–50. 10.1007/s11306-016-1094-6
3.
TurnbaugPLeyRHamadyMFraser-LiggetC. Knight, R, Gordon, J. The human microbiome project. Nature. (2007) 449:804–10. 10.1038/nature06244
4.
RiekebergEPowersR. New frontiers in metabolomics: from measurement to insight. F1000Res. (2017) 6:1148. 10.12688/f1000research.11495.1
5.
SherlockLMokKH. “Metabolomics and Its Applications to Personalized Medicine” in EKC 2019 Conference Proceedings, Springer, Cham (2021), p 25-42.
6.
ZhangASunHYanGWangPWangX. Metabolomics for Biomarker Discovery: Moving to the Clinic. Biomed Res Int. (2015) 2015:e354671. 10.1155/2015/354671
7.
VermeulenRSchymanskiELBarabásiALMillerGW. The exposome and health: where chemistry meets biology. Science. (2020) 367:392–6. 10.1126/science.aay3164
8.
LaurenM. Petrick, Noam S. AI/ML-driven advances in untargeted metabolomics and exposomics for biomedical applications. Cell Rep. Phys. Sci. (2002) 3:2666–3864. 10.1016/j.xcrp.2022.100978
9.
Ter KuileBHWesterhoffHV. Transcriptome meets metabolome hierarchical and metabolic regulation of the glycolytic pathway. FEBS Lett. (2001) 500:169–71. 10.1016/S0014-5793(01)02613-8
10.
GuoLMilburnMVRyalsJALonerganSCMitchellMWWulffJE. Plasma metabolomic profiles enhance precision medicine for volunteers of normal health. Proc Natl Acad Sci USA. (2015) 112:4901–10. 10.1073/pnas.1508425112
11.
NashWJDunnWB. From mass to metabolite in human untargeted metabolomics: Recent advances in annotation of metabolites applying liquid chromatography-mass spectrometry data. TrAC Trends Anal Chem. (2018) 120:e115324. 10.1016/j.trac.2018.11.022
12.
ChongJYamamotoMXiaJ. Metabo Analyst R 2.0: From raw spectra to biological insights. Metabolites. (2019) 9:57–8. 10.3390/metabo9030057
13.
RobertsLDSouzaALGersztenRE. Targeted metabolomics. Curr Protoc Mol Biol. (2012) 30:1–24. 10.1002/0471142727.mb3002s98
14.
DuarteIFDiazSOGilAMNMR. metabolomics of human blood and urine in disease research. J Pharm Biomed Anal. (2014) 93:17–26. 10.1016/j.jpba.2013.09.025
15.
LouisECantrelleFXMesottenLReekmansGBervoetsLVanhoveK. Metabolic phenotyping of human plasma by 1 H-NMR at high and medium magnetic field strengths: a case study for lung cancer. Magn Reson Chem. (2017) 55:706–13. 10.1002/mrc.4577
16.
DettmerKAronovPAHammockBD. Mass spectrometry-based metabolomics. Mass Spec Rev. (2007) 26:51–78. 10.1002/mas.20108
17.
DumezJ-NMilaniJVuichoudBBornetALalande-MartinJTeaI. Hyperpolarized NMR of plant and cancer cell extracts at natural abundance. Analyst. (2015) 140:5860–3. 10.1039/C5AN01203A
18.
LiSParkYDuraisinghamSStrobelFHKhanNSoltowQA. Predicting network activity from high throughput metabolomics. PLOS Comput Biol. (2013) 9:e1003123. 10.1371/journal.pcbi.1003123
19.
PadayacheeTKhamiakovaTLouisEAdriaensensPBurzykowskiT. The impact of the method of extracting metabolic signal from 1H-NMR data on the classification of samples: A case study of binning and BATMAN in lung cancer. PLoS One. (2019) 14:e0211854. 10.1371/journal.pone.0211854
20.
WorleyB. Powers R. Multivariate Analysis in Metabolomics” Curr Metabol. (2013) 1:92–107. 10.2174/2213235X11301010092
21.
MazzellaMSumnerSJGaoSSuLDiaoNMostofaG. Quantitative methods for metabolomic analyses evaluated in the children's health exposure analysis resource (CHEAR). J Expo Sci Environ Epidemiol. (2020) 30:16–27. 10.1038/s41370-019-0162-1
22.
RistMJRothAFrommherzLWeinertCHKruégerRMerzBet al. Metabolite patterns predicting sex and age in participants of the Karlsruhe Metabolomics and Nutrition (KarMeN) study. PLoS ONE. (2017) 12:e0183228. 10.1371/journal.pone.0183228
23.
BubAKriebelADörrCBandtSRistMRothA. The karlsruhe metabolomics and nutrition (KarMeN) study: protocol and methods of a cross-sectional study to characterize the metabolome of healthy men and women. JMIR Res Protoc. (2016) 5:e2603148. 10.2196/resprot.5792
24.
StretchCEastmanTMandalREisnerRWishartDSMourtzakisM. Prediction of skeletal muscle and fat mass in patients with advanced cancer using a metabolomic approach. J Nutr. (2012) 142:14–21. 10.3945/jn.111.147751
25.
FerreaSWintererG. Neuroprotective and neurotoxic effects of nicotine. Pharmacopsychiatry. (2009) 42:255–65. 10.1055/s-0029-1224138
Summary
Keywords
metabolomics, NMR, KarMeN, BATMAN, AI-based algorithm, lung cancer
Citation
Sherlock L, Martin BR, Behsangar S and Mok KH (2023) Application of novel AI-based algorithms to biobank data: uncovering of new features and linear relationships. Front. Med. 10:1162808. doi: 10.3389/fmed.2023.1162808
Received
10 February 2023
Accepted
16 June 2023
Published
13 July 2023
Volume
10 - 2023
Edited by
Stefano Cacciatore, International Centre for Genetic Engineering and Biotechnology (ICGEB), South Africa
Reviewed by
Wimal Pathmasiri, University of North Carolina at Chapel Hill, United States; Cheng-Rong Yu, National Eye Institute (NIH), United States; Yue Victor Zhang, Shenzhen Futian Hospital for Rheumatic Diseases, China
Updates

Check for updates
Copyright
© 2023 Sherlock, Martin, Behsangar and Mok.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: K. H. Mok mok1@tcd.ie
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.