<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neuroinform.</journal-id>
<journal-title>Frontiers in Neuroinformatics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neuroinform.</abbrev-journal-title>
<issn pub-type="epub">1662-5196</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fninf.2020.601829</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Semi-Supervised Learning in Medical Images Through Graph-Embedded Random Forest</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Gu</surname> <given-names>Lin</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/671888/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Xiaowei</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1092129/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>You</surname> <given-names>Shaodi</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhao</surname> <given-names>Shen</given-names></name>
<xref ref-type="aff" rid="aff5"><sup>5</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Liu</surname> <given-names>Zhenzhong</given-names></name>
<xref ref-type="aff" rid="aff6"><sup>6</sup></xref>
<xref ref-type="aff" rid="aff7"><sup>7</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1103951/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Harada</surname> <given-names>Tatsuya</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>RIKEN AIP</institution>, <addr-line>Tokyo</addr-line>, <country>Japan</country></aff>
<aff id="aff2"><sup>2</sup><institution>Research Center for Advanced Science and Technology (RCAST), The University of Tokyo</institution>, <addr-line>Tokyo</addr-line>, <country>Japan</country></aff>
<aff id="aff3"><sup>3</sup><institution>Bioinformatics Institute (BII), A<sup>&#x0002A;</sup>STAR</institution>, <addr-line>Singapore</addr-line>, <country>Singapore</country></aff>
<aff id="aff4"><sup>4</sup><institution>Faculty of Science, Institute of Informatics, University of Amsterdam</institution>, <addr-line>Amsterdam</addr-line>, <country>Netherlands</country></aff>
<aff id="aff5"><sup>5</sup><institution>Department of Medical Physics, Western University</institution>, <addr-line>London, ON</addr-line>, <country>Canada</country></aff>
<aff id="aff6"><sup>6</sup><institution>Tianjin Key Laboratory for Advanced Mechatronic System Design and Intelligent Control, School of Mechanical Engineering, Tianjin University of Technology</institution>, <addr-line>Tianjin</addr-line>, <country>China</country></aff>
<aff id="aff7"><sup>7</sup><institution>National Demonstration Center for Experimental Mechanical and Electrical Engineering Education, Tianjin University of Technology</institution>, <addr-line>Tianjin</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Heye Zhang, Sun Yat-sen University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Chenxi Huang, Xiamen University, China; Guang Yang, Imperial College London, United Kingdom</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Zhenzhong Liu <email>zliu&#x00040;email.tjut.edu.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>10</day>
<month>11</month>
<year>2020</year>
</pub-date>
<pub-date pub-type="collection">
<year>2020</year>
</pub-date>
<volume>14</volume>
<elocation-id>601829</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>09</month>
<year>2020</year>
</date>
<date date-type="accepted">
<day>23</day>
<month>09</month>
<year>2020</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2020 Gu, Zhang, You, Zhao, Liu and Harada.</copyright-statement>
<copyright-year>2020</copyright-year>
<copyright-holder>Gu, Zhang, You, Zhao, Liu and Harada</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p>One major challenge in medical imaging analysis is the lack of label and annotation which usually requires medical knowledge and training. This issue is particularly serious in the brain image analysis such as the analysis of retinal vasculature, which directly reflects the vascular condition of Central Nervous System (CNS). In this paper, we present a novel semi-supervised learning algorithm to boost the performance of random forest under limited labeled data by exploiting the local structure of unlabeled data. We identify the key bottleneck of random forest to be the information gain calculation and replace it with a graph-embedded entropy which is more reliable for insufficient labeled data scenario. By properly modifying the training process of standard random forest, our algorithm significantly improves the performance while preserving the virtue of random forest such as low computational burden and robustness over over-fitting. Our method has shown a superior performance on both medical imaging analysis and machine learning benchmarks.</p></abstract>
<kwd-group>
<kwd>vessel segmentation</kwd>
<kwd>semi-supervised learning</kwd>
<kwd>manifold learning</kwd>
<kwd>central nervous system (CNS)</kwd>
<kwd>retinal image</kwd>
</kwd-group>
<counts>
<fig-count count="4"/>
<table-count count="1"/>
<equation-count count="6"/>
<ref-count count="23"/>
<page-count count="7"/>
<word-count count="3956"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Machine learning has been widely applied to analyze medical images such as an image of the brain. For example, the automatic segmentation of brain tumor (Soltaninejad et al., <xref ref-type="bibr" rid="B19">2018</xref>) could help predict Patient Survival from MRI data. However, traditional methods usually require a large number of diagnosed examples. Collecting raw data during routine screening is possible but making annotations and diagnoses for them is costly and time-consuming for medical experts. To deal with this challenge, we propose a novel graph-embedded semi-supervised algorithm that makes use of the unlabeled data to boost the performance of the random forest. We specifically evaluate the proposed method on both a neuronal image and the retinal image analysis that is highly related to diabetic retinopathy (DR) (Niu et al., <xref ref-type="bibr" rid="B15">2019</xref>) and Alzheimer&#x00027;s Disease (AD) (Liao et al., <xref ref-type="bibr" rid="B12">2018</xref>), and make the following specific contributions:</p>
<list list-type="order">
<list-item><p>We empirically validate that the performance bottleneck of random forest under limited training samples is the biased information gain calculation.</p></list-item>
<list-item><p>We propose a new semi-supervised entropy calculation by incorporating local structure of unlabeled data.</p></list-item>
<list-item><p>We propose a novel semi-supervised random forest which shows advantage performance of the state-of-the-art in both medical imaging analysis and machine learning benchmarks.</p></list-item>
</list>
<p>Among various supervised algorithms, random forest or random decision trees (Breiman et al., <xref ref-type="bibr" rid="B2">1984</xref>; Criminisi et al., <xref ref-type="bibr" rid="B5">2012</xref>) are one of the state-of-the-art machine learning algorithms for medical imaging applications. Despite its robustness and efficiency, its performance relies heavily on sufficiently labeled training data. However, annotating a large amount of medical data is time-consuming and requires domain knowledge. To alleviate the challenge of having enough labeled data, a class of learning methods named semi-supervised learning (SSL) (Joachims, <xref ref-type="bibr" rid="B9">1999</xref>; Zhu et al., <xref ref-type="bibr" rid="B23">2003</xref>; Belkin and Niyogi, <xref ref-type="bibr" rid="B1">2004</xref>; Zhou et al., <xref ref-type="bibr" rid="B21">2004</xref>; Chapelle et al., <xref ref-type="bibr" rid="B4">2006</xref>; Zhu, <xref ref-type="bibr" rid="B22">2006</xref>) were proposed to leverage unlabeled data to improve the performance. Leistner et al. (<xref ref-type="bibr" rid="B10">2009</xref>) proposed a semi-supervised random forest which maximizes the data margin via deterministic annealing (DA). Liu et al. (<xref ref-type="bibr" rid="B14">2015</xref>) showed that the splitting strategy appears to be the bottleneck of performance in a random forest. The authors estimate the unlabeled data through kernel density estimation (KDE) on the projected subspace, and when constructing the internal node, they progressively refine the splitting function with the acquired labels through KDE until it converges. Without explicit affinity relation, CoForest (Li and Zhou, <xref ref-type="bibr" rid="B11">2007</xref>) iteratively guesses the unlabeled data with the rest of the trees in the forest and then uses the new labeled data to refine the tree. Semi-supervised based super-pixel (Gu et al., <xref ref-type="bibr" rid="B7">2017</xref>) has proved to be effective in the segmentation of both a retinal image and a neuronal image.</p>
<p>Following the research line of a previous semi-supervised random forest (RF), we identify that RF&#x00027;s performance bottleneck, under insufficient data, is the biased information gain calculation when selecting an optimal splitting parameter (shown as blue in <xref ref-type="fig" rid="F1">Figure 1</xref>). Therefore, as illustrated in red in <xref ref-type="fig" rid="F1">Figure 1</xref>, we slightly modified the training procedure of RF to relieve this bias. We replace the original information gain with our novel graph-embedded entropy which exploits the data structure of unlabeled data. Specifically, we first use both labeled and unlabeled data to construct a graph whose weights measure local similarity among data and then minimize a loss function that sums the supervised loss over labeled data and a graph Laplacian regularization term. From the optimal solution, we can get label information of unlabeled data which is utilized to estimate a more accurate information gain for node splitting. Since a major part of training and the whole testing remains unchanged, our graph-embedded random forest could significantly improve the performance without losing the virtue of a standard random forest such as low computational burden and robustness over over-fitting.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Difference between our method and standard random forest. Noting that the performance bottleneck (shown in blue) is the biased information gain <italic>G</italic>(&#x003C4;<sub><italic>j</italic></sub>, <italic>w</italic><sub><italic>j</italic></sub>, <italic>X</italic><sub><italic>l</italic></sub>, <italic>Y</italic><sub><italic>l</italic></sub>) calculation based on limited labeled data <italic>X</italic><sub><italic>l</italic></sub>, <italic>Y</italic><sub><italic>l</italic></sub> in Stage 3, we replace <italic>G</italic>(.) with our novel graph-embedded <italic>G</italic><sub><italic>m</italic></sub>(., <italic>X</italic><sub><italic>u</italic></sub>) which considers unlabeled data <italic>X</italic><sub><italic>u</italic></sub> (shown in red).</p></caption>
<graphic xlink:href="fninf-14-601829-g0001.tif"/>
</fig>
</sec>
<sec id="s2">
<title>2. Analysis of Performance Bottleneck</title>
<p>Let us first review the construction of the random forest (Breiman et al., <xref ref-type="bibr" rid="B2">1984</xref>) to figure out why random forest fails under limited training data. A random forest is an ensemble of decision trees: {<italic>t</italic><sub>1</sub>, <italic>t</italic><sub>2</sub>, ..., <italic>t</italic><sub><italic>T</italic></sub>}, of which an individual tree is independently trained and tested.</p>
<p><bold>Training Procedure:</bold> Each decision tree <italic>t</italic>, as illustrated in <xref ref-type="fig" rid="F1">Figure 1</xref>, learns to classify a training sample <italic>x</italic> &#x02208; X to the corresponding label <italic>y</italic> by recursively branching it to the left or right child until reaching a leaf node. In particular, each node is associated with a binary split function <italic>h</italic>(<italic>x</italic><sub><italic>i</italic></sub>, <italic>w</italic>, &#x003C4;), e.g., oblique linear split function</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>h</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow><mml:mo>&#x0003C;</mml:mo><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where [.] is an indicative function and &#x003C4; is a scaler threshold. <italic>w</italic> &#x02208; <italic>R</italic><sup><italic>d</italic></sup> serves as a feature weight parameter that projects the high dimension data <italic>x</italic> &#x02208; <italic>R</italic><sup><italic>d</italic></sup> to a one dimensional subspace.</p>
<p>Given a candidate splitting function <italic>h</italic>(<italic>x, w</italic><sub><italic>j</italic></sub>, &#x003C4;<sub><italic>j</italic></sub>), its splitting quality is measured by information gain <italic>G</italic>(<italic>w</italic><sub><italic>j</italic></sub>, &#x003C4;<sub><italic>j</italic></sub>). In practice, given the training data <italic>X</italic> and their labels <italic>Y</italic>, the construction of the splitting node, as illustrated in the left side of <xref ref-type="fig" rid="F1">Figure 1</xref>, comprises the following three stages:</p>
<table-wrap position="float">
<label>Algorithm 1</label>
<caption><p>Training of node splitting.</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr>
<td align="left" valign="top">1: &#x000A0;Randomly generates a set of feature subspace candidates {<italic>w</italic><sub><italic>j</italic></sub>}</td>
</tr>
<tr>
<td align="left" valign="top">2: &#x000A0;For each <italic>w</italic><sub><italic>j</italic></sub>, find the optimal <inline-formula><mml:math id="M2"><mml:msubsup><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mtext>argmax</mml:mtext></mml:mrow><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow></mml:munder><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x003C4;</mml:mi><mml:mo>,</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> that best splits the data.</td>
</tr>
<tr>
<td align="left" valign="top">3: &#x000A0;Among all {<italic>w</italic><sub><italic>j</italic></sub>, <inline-formula><mml:math id="M3"><mml:msubsup><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>}, pick the one with largest information gain: <inline-formula><mml:math id="M4"><mml:msup><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mtext>argmax</mml:mtext></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula></td>
</tr> 
</tbody>
</table>
</table-wrap>
<p>Through the above stages, each split node is associated with a splitting function <italic>h</italic>(<italic>x, w</italic>, &#x003C4;) that best splits the training data.</p>
<p><bold>Testing Procedure:</bold> When testing data <italic>x</italic>, the trained random forest predicts the probability of its label by averaging the ensemble prediction as <inline-formula><mml:math id="M5"><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where <italic>p</italic><sub><italic>t</italic></sub>(<italic>y</italic>|<italic>x</italic>) denotes the empirical label distribution of the training samples that reach leaf note of tree <italic>t</italic>.</p>
<sec>
<title>2.1. Performance Bottleneck Under Insufficient Data</title>
<p>According to the study of Liu et al. (<xref ref-type="bibr" rid="B14">2015</xref>), insufficient training data would impact the performance of RF in three ways (Liu et al., <xref ref-type="bibr" rid="B14">2015</xref>): (1) limited forest depth; (2) inaccurate prediction model of leaf nodes; (3) sub-optimal splitting strategy. Among them, Liu et al. (<xref ref-type="bibr" rid="B14">2015</xref>) identified that (1) is inevitable, and (2) is solvable with their proposed strategy. In this paper, we further improve the method by tackling (3).</p>
<p>We claim that the performance bottleneck of random forest is its sub-optimal splitting strategy in Algorithm 1. To empirically support this claim, we build three random forests, similar to Liu et al. (<xref ref-type="bibr" rid="B14">2015</xref>), for comparison: the first one, the <bold>Control</bold> is trained with a small size of a training set <italic>S</italic>1 as control; the second one, the <bold>Perfect Stage 3</bold> is constructed with the same training set <italic>S</italic>1 but its node splitting uses a large training set <italic>S</italic>2 to select the optimal parameter in stage 3 of Algorithm 1, to simulate the case that random forest selects the optimal parameter of stage 3 with full information; the third one, the <bold>Perfect Splitting</bold> is constructed with <italic>S</italic>1 while <italic>S</italic>2 was used for both Stage 2 and 3 of Algorithm 1.</p>
<p>Following the protocol of Liu et al. (<xref ref-type="bibr" rid="B14">2015</xref>), each random forest comprises 100 trees and the same entropy gain is adopted as the splitting criterion. We evaluate three random forests on Madelon (Guyon et al., <xref ref-type="bibr" rid="B8">2004</xref>), a widely used machine learning benchmark. As shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, Perfect Stage 3, which only uses the full information to select the best parameter set, significantly improves the performance compared to the control group. Interestingly, the Perfect Splitting one, which utilizes the full information for both optimal parameter proposing (Stage 2) and optimal parameter decision (Stage 3), only makes a subtle improvement compared to Perfect Stage 3.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Empirical validation of performance Bottleneck.</p></caption>
<graphic xlink:href="fninf-14-601829-g0002.tif"/>
</fig>
<p>From <xref ref-type="fig" rid="F2">Figure 2</xref>, we found that Stage 3, optimal parameter selection, is the performance bottleneck of the splitting node construction, which is also the keystone of random forest construction (Liu et al., <xref ref-type="bibr" rid="B14">2015</xref>). When deciding the optimal parameter, random forest often fails to find the best one as its information gain calculation <italic>g</italic>(<italic>w</italic>, &#x003C4;) is biased under insufficient training data. Interestingly, insufficient data has a smaller effect on the Stage 2, parameter proposal. Motivated by this observation, we propose a new information calculation which exploits unlabeled data to make a better parameter selection in Stage 3 of Algorithm 1.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Graph-Embedded Representation of Information Gain</title>
<p>In the previous section, we show that gain estimation appears to be the performance bottleneck of random forests. Empirically, we show that more label information helps to obtain more accurate gain estimation. This encourages us to consider the possibility of mining label information from unlabeled data through structural connections between labeled and unlabeled data. In particular, we perform a graph-based semi-supervised learning to get label information of unlabeled data, and compute information gain from both labeled and unlabeled data. To achieve a better gain estimation, we embed all data into a graph. Moreover, we assume the underlying structure of all data form a manifold, and compute data similarity based on the assumption.</p>
<p>Let <italic>l</italic> and <italic>u</italic> be the number of labeled and unlabeled instances, respectively. Let <inline-formula><mml:math id="M6"><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mspace width="0.3em" class="thinspace"/><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> be the matrix of feature vectors of labeled instances, and <inline-formula><mml:math id="M7"><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mspace width="0.3em" class="thinspace"/><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>&#x022A4;</mml:mo></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>u</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> be the matrix of unlabeled instances. To accommodate label information, we define a label matrix <italic>Y</italic> &#x02208; &#x0211D;<sup>(<italic>l</italic>&#x0002B;<italic>u</italic>) &#x000D7; <italic>K</italic></sup> (assuming there are <italic>K</italic> class labels available), with each entry <italic>Y</italic><sub><italic>ik</italic></sub> containing 1 provided the <italic>i</italic>-th data belongs to <italic>X</italic><sub><italic>l</italic></sub> and is labeled with class <italic>k</italic>, and 0 otherwise. Besides, we define <italic>Y</italic><sub><italic>l</italic></sub> as a submatrix of <italic>Y</italic> corresponding to the labeled data, <inline-formula><mml:math id="M8"><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> as the <italic>i</italic>-th row of <italic>Y</italic> corresponding to <italic>x</italic><sub><italic>l</italic></sub>, and <inline-formula><mml:math id="M9"><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>y</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> as the vector of class labels for <italic>X</italic><sub><italic>l</italic></sub>.</p>
<p>Based on both labeled and unlabeled instances, our purpose is to learn a mapping <italic>f</italic> : &#x0211D;<sup><italic>d</italic></sup> &#x02192; &#x0211D;<sup><italic>K</italic></sup> and predict the label of instance <italic>x</italic> as <inline-formula><mml:math id="M10"><mml:msup><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msup><mml:mo>:</mml:mo><mml:mo>=</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">arg</mml:mtext></mml:mstyle><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. Many semi-supervised learning algorithms use the following regularized framework</p>
<disp-formula id="E2"><mml:math id="M11"><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mtext class="textit" mathvariant="italic">loss</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>u</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msubsup><mml:mrow><mml:mo stretchy="false">&#x02016;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo stretchy="false">&#x02016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>loss</italic>() is a loss function and <italic>s</italic>(<italic>x</italic><sub><italic>i</italic></sub>, <italic>x</italic><sub><italic>j</italic></sub>) is a similarity function. In this paper, we apply the idea of graph embedding to learn <italic>f</italic>. We construct a graph <inline-formula><mml:math id="M12"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">G</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>E</mml:mi><mml:mo>,</mml:mo><mml:mi>W</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where each node in <italic>V</italic> denotes a training instance and <italic>W</italic> &#x02208; &#x0211D;<sup>(<italic>l</italic>&#x0002B;<italic>u</italic>) &#x000D7; (<italic>l</italic>&#x0002B;<italic>u</italic>)</sup> denotes a symmetric weight matrix. <italic>W</italic> is computed as follows: for each point find <italic>t</italic> nearest neighbors, and <inline-formula><mml:math id="M13"><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mo stretchy="false">&#x02016;</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">&#x02016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>/</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> if (<italic>x</italic><sub><italic>i</italic></sub>, <italic>x</italic><sub><italic>j</italic></sub>) are neighbors, 0 otherwise. Such construction of graph implicitly assumes that all data resides on some manifold and exploits local structure. Based on the graph embedding, we propose to minimize</p>
<disp-formula id="E3"><label>(2)</label><mml:math id="M14"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>u</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mo stretchy="false">&#x02016;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mrow><mml:mo stretchy="false">&#x02016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>u</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:mo stretchy="false">|</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac><mml:mo stretchy="false">|</mml:mo><mml:msubsup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>D</italic> is a diagonal matrix with its <italic>D</italic><sub><italic>ii</italic></sub> equal to the sum of the <italic>i</italic>-th row of <italic>W</italic>. Let <inline-formula><mml:math id="M15"><mml:msup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mspace width="0.3em" class="thinspace"/><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">arg</mml:mtext></mml:mstyle><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:munder><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> be the optimal solution, it has been shown in Zhou et al. (<xref ref-type="bibr" rid="B21">2004</xref>) that</p>
<disp-formula id="E4"><label>(3)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>I</mml:mi><mml:mo>-</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:msup><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mi>W</mml:mi><mml:msup><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mi>Y</mml:mi><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Based on the learned functions <italic>F</italic><sup>&#x0002A;</sup>, we can predict the label information of <italic>X</italic><sub><italic>u</italic></sub> and then utilize such information to estimate more accurate information gain. Specifically, we let <inline-formula><mml:math id="M17"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>y</mml:mi></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the predicted label of <italic>X</italic><sub><italic>u</italic></sub>, and for node <italic>S</italic> we compute Gini index <inline-formula><mml:math id="M18"><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, where</p>
<disp-formula id="E5"><mml:math id="M19"><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x02264;</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mo>&#x1D7D9;</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>y</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x02264;</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>u</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mo>&#x1D7D9;</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>y</mml:mi></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>is the proportion of data from class <italic>k</italic>. Note that we utilize information from both labeled and unlabeled data to compute the Gini index. For each node, we estimate information gain as</p>
<disp-formula id="E6"><label>(4)</label><mml:math id="M20"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003C4;</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mo stretchy="false">|</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>S</italic><sub><italic>l</italic></sub> and <italic>S</italic><sub><italic>u</italic></sub> are left and right child nodes, respectively.</p>
</sec>
<sec id="s4">
<title>4. Construction of Semi-Supervised Random Forest</title>
<p>In our framework, we preserve the major structure of the standard random forest where the testing stage is exactly the same as the standard one. As illustrated in the right part of <xref ref-type="fig" rid="F1">Figure 1</xref>, we only make a small modification in stage 3 of Algorithm 1 where the splitting efficiency is now evaluated by our novel graph-embedded based information gain <italic>G</italic><sub><italic>m</italic></sub>(&#x003C4;<sub><italic>j</italic></sub>, <italic>w</italic><sub><italic>j</italic></sub>, <italic>X</italic><sub><italic>l</italic></sub>, <italic>Y</italic><sub><italic>l</italic></sub>, <italic>X</italic><sub><italic>u</italic></sub>) from Equation (4). Specifically, we leave stage 2 unchanged that the threshold &#x003C4; of each subspace candidate <italic>w</italic> is still based on standard information gain such as the Gini index. Now with a set of parameter candidates <italic>w</italic>, &#x003C4;, the stage 3 calculates the corresponding manifold based information score &#x0011D;(<italic>w</italic>, &#x003C4;) instead and select the optimal one through <inline-formula><mml:math id="M21"><mml:munder><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">max</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mi>&#x0011D;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
</sec>
<sec id="s5">
<title>5. Experiments</title>
<p>We evaluate our method on both 2D, and 3D brain related medical image segmentation tasks as well as two machine learning benchmarks.</p>
<p>The retinal vessel, a part of the Central Nervous System (CNS), directly reflects the vascular condition of CNS. The accurate segmentation of vessels is important for this analysis. Much progress has been made based on either random forest (Gu et al., <xref ref-type="bibr" rid="B6">2017</xref>) or deep learning (Liu et al., <xref ref-type="bibr" rid="B13">2019</xref>). The DRIVE dataset (Staal et al., <xref ref-type="bibr" rid="B20">2004</xref>) is a widely used 2D retinal vessel segmentation dataset that comprises of 20 training images and 20 testing ones. Each image is a 768 &#x000D7;584 color image along with manual segmentation. For the image, we extract two types of widely used features: 1, local patch <inline-formula><mml:math id="M22"><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mn>15</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>15</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> of target. 2, <inline-formula><mml:math id="M23"><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>7</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> Gabor wavelets (Soares et al., <xref ref-type="bibr" rid="B18">2006</xref>). We also investigate the single neuron segmentation in a brain image. BigNeuron project<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> (Peng et al., <xref ref-type="bibr" rid="B16">2015</xref>) is a 3D neuronal dataset with ground truth annotation from experts. For BigNeuron data, we manually picked 13 images among which a random 10 were used for training while the rest were left for testing, because this dataset is designed for tracing rather than segmentation. For example, some annotation is visibly thinner than the actual neuron. Furthermore, the image may contain multiple neurons but only one is properly annotated. For both datasets, we randomly collected 40, 000 (20, 000 positive and 20, 000 negative) samples from the training and testing sets, respectively. For 3D data, our feature is <inline-formula><mml:math id="M24"><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mn>15</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>15</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>7</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> local cube similar to the setting of Gu et al. (<xref ref-type="bibr" rid="B6">2017</xref>).</p>
<p>Apart from the medical imaging, we also demonstrate the generality of our method on two binary machine learning benchmark, IJCNN1 (Prokhorov, <xref ref-type="bibr" rid="B17">2001</xref>) and Madelon (Guyon et al., <xref ref-type="bibr" rid="B8">2004</xref>), in Libsvm Repository (Chang and Lin, <xref ref-type="bibr" rid="B3">2011</xref>).</p>
<p>During the evaluation, we randomly selected a certain number <italic>n</italic> of labeled samples from the whole training set while leaving the rest unlabeled. Standard Random Forest (RF) is trained with <italic>n</italic> labeled training data only. Our method and RobustNode (Liu et al., <xref ref-type="bibr" rid="B14">2015</xref>) are trained with both labeled data and unlabeled data. For reference, we also compared it with Optimal RF which is trained with labeled data as a standard RF. However, its node splitting is supervised with the whole training samples and their label. Optimal RF indicates the upper bound for all of semi-supervised learning algorithms.</p>
<sec>
<title>5.1. Medical Imaging Segmentation</title>
<p>First, we illustrate the visual performance of segmentation in <xref ref-type="fig" rid="F3">Figure 3</xref>. The estimated score is the possibility of the vessel given by the individual method. Our algorithm has consistently improved the estimation compared to the standard RF.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Exemplar estimation of vessel on the DRIVE dataset with 800 labeled samples. From left to right: Input images; Ground-truth; Estimation of our method; Estimation of Standard RF; Estimation of Optimal RF.</p></caption>
<graphic xlink:href="fninf-14-601829-g0003.tif"/>
</fig>
</sec>
<sec>
<title>5.2. Quantitative Analysis</title>
<p>We also report the classification accuracy with respect to the number of labeled data in <xref ref-type="fig" rid="F4">Figure 4</xref>, <xref ref-type="table" rid="T1">Table 1</xref>. We compared our method with alternatives on both medical imaging segmentation and machine learning benchmarks. <xref ref-type="fig" rid="F4">Figure 4</xref> shows that our algorithm significantly outperformed alternative methods. Specifically, in the DRIVE dataset, our algorithm approaches the upper bound at 1,000 labeled samples. In the IJCNN1 dataset, our method quickly approaches the optimal one while the alternatives take 400 samples to approach.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Classification accuracy vs. number of labeled samples.</p></caption>
<graphic xlink:href="fninf-14-601829-g0004.tif"/>
</fig>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Classification accuracy (represented in percentage %) on different dataset.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th/>
<th valign="top" align="center"><bold>Drive</bold></th>
<th valign="top" align="center"><bold>Big neuron</bold></th>
<th valign="top" align="center"><bold>IJCNN1</bold></th>
<th valign="top" align="center"><bold>Madelon</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Our method</td>
<td valign="top" align="center">79.42</td>
<td valign="top" align="center">74.16</td>
<td valign="top" align="center">89.36</td>
<td valign="top" align="center">59.57</td>
</tr>
<tr>
<td valign="top" align="left">Standard RF</td>
<td valign="top" align="center">60.90</td>
<td valign="top" align="center">70.93</td>
<td valign="top" align="center">79.89</td>
<td valign="top" align="center">51.33</td>
</tr>
<tr>
<td valign="top" align="left">Robust node RF</td>
<td valign="top" align="center">63.79</td>
<td valign="top" align="center">70.99</td>
<td valign="top" align="center">78.92</td>
<td valign="top" align="center">50.53</td>
</tr>
<tr>
<td valign="top" align="left">Optimal RF</td>
<td valign="top" align="center">85.53</td>
<td valign="top" align="center">75.55</td>
<td valign="top" align="center">91.29</td>
<td valign="top" align="center">67.10</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>We show the accuracy on the training sample of 400 (DRIVE), 1,500 (Big Neuron), 300 (IJCNN1), and 400 (Madelon)</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec sec-type="conclusions" id="s6">
<title>6. Conclusion</title>
<p>In this paper, we propose a novel semi-supervised random forest to tackle the challenging problem of the lacking annotation in the analysis of medical imaging such as a brain image. Observing that the bottleneck of the standard random forest is the biased information gain estimation, we replaced it with a novel graph-embedded entropy which incorporates information from both labeled and unlabeled data. Empirical results show that our information gain is more reliable than the one used in traditional random forest under insufficient labeled data. By slightly modifying the training process of the standard random forest, our algorithm significantly improves the performance while preserving the virtue of the random forest. Our method has shown a superior performance with very limited data in both brain imaging analysis and machine learning benchmarks.</p>
</sec>
<sec sec-type="data-availability-statement" id="s7">
<title>Data Availability Statement</title>
<p>All datasets generated for this study are included in the article/supplementary material.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>All authors listed have made a substantial, direct and intellectual contribution to the work, and approved it for publication.</p>
</sec>
<sec id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Belkin</surname> <given-names>M.</given-names></name> <name><surname>Niyogi</surname> <given-names>P.</given-names></name></person-group> (<year>2004</year>). <article-title>Semi-supervised learning on Riemannian manifolds</article-title>. <source>Mach. Learn</source>. <volume>56</volume>, <fpage>209</fpage>&#x02013;<lpage>239</lpage>. <pub-id pub-id-type="doi">10.1023/B:MACH.0000033120.25363.1e</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Breiman</surname> <given-names>L.</given-names></name> <name><surname>Friedman</surname> <given-names>J.</given-names></name> <name><surname>Stone</surname> <given-names>C. J.</given-names></name> <name><surname>Olshen</surname> <given-names>R. A.</given-names></name></person-group> (<year>1984</year>). <source>Classification And Regression Trees</source>. <publisher-name>CRC Press</publisher-name>.</citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chang</surname> <given-names>C.-C.</given-names></name> <name><surname>Lin</surname> <given-names>C.-J.</given-names></name></person-group> (<year>2011</year>). <article-title>Libsvm: A library for support vector machines</article-title>. <source>ACM Trans. Intell. Syst. Technol</source>. <volume>2</volume>:<fpage>27</fpage>. <pub-id pub-id-type="doi">10.1145/1961189.1961199</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chapelle</surname> <given-names>O.</given-names></name> <name><surname>Scholkopf</surname> <given-names>B.</given-names></name> <name><surname>Zien</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <source>Semi-Supervised Learning</source>. <publisher-loc>London</publisher-loc>: <publisher-name>MIT Press</publisher-name>. <pub-id pub-id-type="doi">10.7551/mitpress/9780262033589.001.0001</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Criminisi</surname> <given-names>A.</given-names></name> <name><surname>Shotton</surname> <given-names>J.</given-names></name> <name><surname>Konukoglu</surname> <given-names>E.</given-names></name></person-group> (<year>2012</year>). <article-title>Decision forests: a unified framework for classification, regression, density estimation, manifold learning and semi-supervised learning</article-title>. <source>Found. Trends. Comput. Graph. Vis</source>. <volume>7</volume>, <fpage>81</fpage>&#x02013;<lpage>227</lpage>. <pub-id pub-id-type="doi">10.1561/0600000035</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Cheng</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>Segment 2D and 3D filaments by learning structured and contextual features</article-title>. <source>IEEE Trans. Med. Imaging</source> <volume>36</volume>, <fpage>596</fpage>&#x02013;<lpage>606</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2016.2623357</pub-id><pub-id pub-id-type="pmid">27831862</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>L.</given-names></name> <name><surname>Zheng</surname> <given-names>Y.</given-names></name> <name><surname>Bise</surname> <given-names>R.</given-names></name> <name><surname>Sato</surname> <given-names>I.</given-names></name> <name><surname>Imanishi</surname> <given-names>N.</given-names></name> <name><surname>Aiso</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>Semi-supervised learning for biomedical image segmentation via forest oriented super pixels(voxels)</article-title>, in <source>Medical Image Computing and Computer Assisted Intervention MICCAI 2017</source>, eds <person-group person-group-type="editor"><name><surname>Descoteaux</surname> <given-names>M.</given-names></name> <name><surname>Maier-Hein</surname> <given-names>L.</given-names></name> <name><surname>Franz</surname> <given-names>A.</given-names></name> <name><surname>Jannin</surname> <given-names>P.</given-names></name> <name><surname>Collins</surname> <given-names>D. L.</given-names></name> <name><surname>Duchesne</surname> <given-names>S.</given-names></name></person-group> (<publisher-loc>Quebec City, QC</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>702</fpage>&#x02013;<lpage>710</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-66182-7_80</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Guyon</surname> <given-names>I.</given-names></name> <name><surname>Gunn</surname> <given-names>S.</given-names></name> <name><surname>Hur</surname> <given-names>A. B.</given-names></name> <name><surname>Dror</surname> <given-names>G.</given-names></name></person-group> (<year>2004</year>). <article-title>Result analysis of the NIPS 2003 feature selection challenge</article-title>, in <source>Proceedings of the 17th International Conference on Neural Information Processing Systems, NIPS&#x00027;04</source> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>545</fpage>&#x02013;<lpage>552</lpage>.</citation></ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Joachims</surname> <given-names>T.</given-names></name></person-group> (<year>1999</year>). <article-title>Transductive inference for text classification using support vector machines</article-title>, in <source>Proceedings of the Sixteenth International Conference on Machine Learning, ICML &#x00027;99</source> (<publisher-loc>San Francisco, CA</publisher-loc>: <publisher-name>Morgan Kaufmann Publishers Inc.</publisher-name>), <fpage>200</fpage>&#x02013;<lpage>209</lpage>. <pub-id pub-id-type="pmid">19623491</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Leistner</surname> <given-names>C.</given-names></name> <name><surname>Saffari</surname> <given-names>A.</given-names></name> <name><surname>Santner</surname> <given-names>J.</given-names></name> <name><surname>Bischof</surname> <given-names>H.</given-names></name></person-group> (<year>2009</year>). <article-title>Semi-supervised random forests</article-title>, in <source>2009 IEEE 12th International Conference on Computer Vision</source> (<publisher-loc>Kyoto</publisher-loc>), <fpage>506</fpage>&#x02013;<lpage>513</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2009.5459198</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Zhou</surname> <given-names>Z. H.</given-names></name></person-group> (<year>2007</year>). <article-title>Improve computer-aided diagnosis with machine learning techniques using undiagnosed samples</article-title>. <source>IEEE Trans. Syst. Man Cybern</source>. <volume>37</volume>, <fpage>1088</fpage>&#x02013;<lpage>1098</lpage>. <pub-id pub-id-type="doi">10.1109/TSMCA.2007.904745</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liao</surname> <given-names>H.</given-names></name> <name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Peng</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>Potential utility of retinal imaging for Alzheimers disease: a review</article-title>. <source>Front. Aging Neurosci</source>. <volume>10</volume>:<fpage>188</fpage>. <pub-id pub-id-type="doi">10.3389/fnagi.2018.00188</pub-id><pub-id pub-id-type="pmid">29988470</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Gu</surname> <given-names>L.</given-names></name> <name><surname>Lu</surname> <given-names>F.</given-names></name></person-group> (<year>2019</year>). <article-title>Unsupervised ensemble strategy for retinal vessel segmentation</article-title>, in <source>Medical Image Computing and Computer Assisted Intervention-MICCAI 2019</source>, eds <person-group person-group-type="editor"><name><surname>Shen</surname> <given-names>D.</given-names></name> <name><surname>Liu</surname> <given-names>T.</given-names></name> <name><surname>Peters</surname> <given-names>T. M.</given-names></name> <name><surname>Staib</surname> <given-names>L. H.</given-names></name> <name><surname>Essert</surname> <given-names>C.</given-names></name> <name><surname>Zhou</surname> <given-names>S.</given-names></name> <name><surname>Yap</surname> <given-names>P. T.</given-names></name> <name><surname>Khan</surname> <given-names>A.</given-names></name></person-group> (<publisher-loc>Shenzhen</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>111</fpage>&#x02013;<lpage>119</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-32239-7_13</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Song</surname> <given-names>M.</given-names></name> <name><surname>Tao</surname> <given-names>D.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Chen</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Random forest construction with robust semisupervised node splitting</article-title>. <source>IEEE Trans. Image Process</source>. <volume>24</volume>, <fpage>471</fpage>&#x02013;<lpage>483</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2014.2378017</pub-id><pub-id pub-id-type="pmid">25494503</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Niu</surname> <given-names>Y.</given-names></name> <name><surname>Gu</surname> <given-names>L.</given-names></name> <name><surname>Lu</surname> <given-names>F.</given-names></name> <name><surname>Lv</surname> <given-names>F.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Sato</surname> <given-names>I.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Pathological evidence exploration in deep retinal image diagnosis</article-title>, in <source>AAAI conference on artificial intelligence (AAAI)</source> (<publisher-loc>Honolulu</publisher-loc>: <publisher-name>AAAI Press</publisher-name>), <fpage>1093</fpage>&#x02013;<lpage>1101</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v33i01.33011093</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peng</surname> <given-names>H.</given-names></name> <name><surname>Hawrylycz</surname> <given-names>M.</given-names></name> <name><surname>Roskams</surname> <given-names>J.</given-names></name> <name><surname>Hill</surname> <given-names>S.</given-names></name> <name><surname>Spruston</surname> <given-names>N.</given-names></name> <name><surname>Meijering</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Bigneuron: large-scale 3D neuron reconstruction from optical microscopy images</article-title>. <source>Neuron</source> <volume>87</volume>, <fpage>252</fpage>&#x02013;<lpage>256</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2015.06.036</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Prokhorov</surname> <given-names>D.</given-names></name></person-group> (<year>2001</year>). <source>IJCNN 2001 Neural Network Competition</source>. <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>IJCNN2001</publisher-name>.</citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soares</surname> <given-names>J. V. B.</given-names></name> <name><surname>Leandro</surname> <given-names>J. J. G.</given-names></name> <name><surname>Cesar</surname> <given-names>R. M.</given-names></name> <name><surname>Jelinek</surname> <given-names>H. F.</given-names></name> <name><surname>Cree</surname> <given-names>M. J.</given-names></name></person-group> (<year>2006</year>). <article-title>Retinal vessel segmentation using the 2-D gabor wavelet and supervised classification</article-title>. <source>IEEE Trans. Med. Imaging</source> <volume>25</volume>, <fpage>1214</fpage>&#x02013;<lpage>1222</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2006.879967</pub-id><pub-id pub-id-type="pmid">16967806</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Soltaninejad</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Lambrou</surname> <given-names>T.</given-names></name> <name><surname>Yang</surname> <given-names>G.</given-names></name> <name><surname>Allinson</surname> <given-names>N.</given-names></name> <name><surname>Ye</surname> <given-names>X.</given-names></name></person-group> (<year>2018</year>). <article-title>MRI brain tumor segmentation and patient survival prediction using random forests and fully convolutional networks</article-title>, in <source>Brainlesion: Glioma, Multiple Sclerosis, Stroke and Traumatic Brain Injuries</source>, eds <person-group person-group-type="editor"><name><surname>Crimi</surname> <given-names>A.</given-names></name> <name><surname>Bakas</surname> <given-names>S.</given-names></name> <name><surname>Kuijf</surname> <given-names>H.</given-names></name><name><surname>Menze</surname> <given-names>B.</given-names></name> <name><surname>Reyes</surname> <given-names>M.</given-names></name></person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>204</fpage>&#x02013;<lpage>215</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-75238-9_18</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Staal</surname> <given-names>J.</given-names></name> <name><surname>Abramoff</surname> <given-names>M. D.</given-names></name> <name><surname>Niemeijer</surname> <given-names>M.</given-names></name> <name><surname>Viergever</surname> <given-names>M. A.</given-names></name> <name><surname>van Ginneken</surname> <given-names>B.</given-names></name></person-group> (<year>2004</year>). <article-title>Ridge-based vessel segmentation in color images of the retina</article-title>. <source>IEEE Trans. Med. Imaging</source> <volume>23</volume>, <fpage>501</fpage>&#x02013;<lpage>509</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2004.825627</pub-id><pub-id pub-id-type="pmid">15084075</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>D.</given-names></name> <name><surname>Bousquet</surname> <given-names>O.</given-names></name> <name><surname>Lal</surname> <given-names>T.</given-names></name> <name><surname>Weston</surname> <given-names>J.</given-names></name> <name><surname>Scholkopf</surname> <given-names>B.</given-names></name></person-group> (<year>2004</year>). <article-title>Learning with local and global consistency</article-title>, in <source>NIPS</source> (<publisher-loc>Vancouver, BC</publisher-loc>).</citation></ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>X.</given-names></name></person-group> (<year>2006</year>). <source>Semi-Supervised Learning Literature Survey</source>. Technical report, University of Wisconsin-Madison.</citation></ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>X.</given-names></name> <name><surname>Ghahramani</surname> <given-names>Z.</given-names></name> <name><surname>Lafferty</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2003</year>). <article-title>Semi-supervised learning using Gaussian fields and harmonic functions</article-title>, in <source>ICML, Vol. 3</source> (<publisher-loc>Washington, DC</publisher-loc>), <fpage>912</fpage>&#x02013;<lpage>919</lpage>.</citation></ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://www.alleninstitute.org/bigneuron/about/">https://www.alleninstitute.org/bigneuron/about/</ext-link></p></fn>
</fn-group>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> This research was supported by JST, ACT-X Grant Number JPMJAX190D, Japan and the National Natural Science Foundation of China (Grant No. 61873188).</p>
</fn>
</fn-group>
</back>
</article>