ORIGINAL RESEARCH article

Front. Plant Sci., 01 July 2021

Sec. Technical Advances in Plant Science

Volume 12 - 2021 | https://doi.org/10.3389/fpls.2021.630425

TheLNet270v1 – A Novel Deep-Network Architecture for the Automatic Classification of Thermal Images for Greenhouse Plants

  • 1. Agricultural AI Research Promotion Office, RCAIT, National Agriculture and Food Research Organization (NARO), Tsukuba, Japan

  • 2. Institute of Vegetable and Flower Research, NARO, Tsukuba, Japan

Abstract

The real challenge for separating leaf pixels from background pixels in thermal images is associated with various factors such as the amount of emitted and reflected thermal radiation from the targeted plant, absorption of reflected radiation by the humidity of the greenhouse, and the outside environment. We proposed TheLNet270v1 (thermal leaf network with 270 layers version 1) to recover the leaf canopy from its background in real time with higher accuracy than previous systems. The proposed network had an accuracy of 91% (mean boundary F1 score or BF score) to distinguish canopy pixels from background pixels and then segment the image into two classes: leaf and background. We evaluated the classification (segment) performance by using more than 13,766 images and obtained 95.75% training and 95.23% validation accuracies without overfitting issues. This research aimed to develop a deep learning technique for the automatic segmentation of thermal images to continuously monitor the canopy surface temperature inside a greenhouse.

Introduction

Leaf surface and internal structure changes are due to adverse growth, stomatal resistance, diseases, leaf angles, depth of the canopy, and water stress conditions, which alter the absorbance-reflection process of solar radiation (; ; ). Thermography detected this reflected (emitted) long-wave infrared (8–14 μm), then converted it into thermal images, and a false-color gradient demonstrated the temperature level of the plant leaves of canopies (). Figure 1 shows the working principle of a thermal camera.

Figure 1

Over the last few years, the advancement of fast computing power, low-cost imaging systems with image processing software, and deep learning (DL) techniques have allowed for nondestructive disease diagnosis and detection of various stress conditions of plants in a timely manner (). The DL based on a convolution neural network (CNN) is the successor of traditional machine learning approaches that can learn features with greater precision and accuracy by activating maximum networkability (). compared CNN-based DL with the Neocortex of the human brain, which learns response-based features dynamically from images. CNN-based DL acquires hierarchical features and emphasizes nonlinear filters of the depth of the deep network structure for learning, and after that solves problem-specific tasks such as image classification, semantic segmentation (pixel-based classification), object detection, video processing, speech recognition, and natural language processing (Simonyan and Zisserman, 2014; ; ). classified deep network architectures into seven classes: spatial exploitation, depth, multi-path, width, feature-map exploitation, channel boosting, and attention-based CNNs. Figure 2 demonstrates the classification of various deep network architectures along with the proposed TheLNet270v1. stated that a filter termed as a channel in a CNN can extract different levels of information (from fine-grained to coarse-grained) based on their sizes (small to large sizes). and stated that the deep DL architecture has an advantage over the shallow depth DL architecture, which can learn complex representations at different levels of abstraction and thus increase the classification accuracy. According to , branching within layers can abstract features with various spatial scales. , , , , , , , and proposed multi-paths or shortcut connections that connect one layer with another layer by skipping some intermediate layers. This allows overpassing some information to another layer and reduces the vanishing gradient problem, which causes a higher training error.

Figure 2

proposed an edge-conditioned convolution neural network for thermal image segmentation with SODA (segment objects in day and night) benchmark for evaluating the thermal image segmentation performance. They used manually annotated synthetically generated thermal images for training the network, which achieved 61.9% mean intersection over union (IoU), a lightly better than network trained with DeepLabv3 algorithm. developed a CNN-based thermal Image Enhancement technique for improving low-resolution thermal camera recognition tasks. The lightweight structure of the shallow convolutional neural network requires less CPU memory. In their architecture, they cropped a low-resolution thermal image with a uniform stride and used a bi-cubic interpolation method to upscale it. revealed a Fletcher-Reeves algorithm-based CNN model for hyperspectral image classification with 80.7% accuracy which outperforms other traditional CNN due to the advances in batch computing adaptability and convergence speed. reported a thermal-RGB image-based wheat-ears detection system for automatic counting wheat wars under outdoor conditions. They applied blocks of convolutional layers, each with an activated function for counting the wheat ears, which achieved 75.63 and 68.46% F1 Score for segmenting thermal and RGB images. Furthermore, achieved 89.22% accuracy for counting the wheat ears. Another study reported by used a group of neurons termed as a capsule or vector for replacing traditional neurons and achieved equivariance by successfully encoding spatial information and properties of an input image. identified and classified objects in real time from thermal cameras carried by firefighters. The detection accuracy reported by authors varied from 70 to 95%, which depends on the depth of the convolution network layer.

developed a CNN with an encoder–decoder function which used top-view RGB images of fig plants and achieved a mean 93.85% segmentation accuracy under variable visual fig leaves the appearance and complex background. There are various DL architectures, such as LeNet, AlexNet, VGG, GoogleNet, YOLOv, Inception, and SqueezeNet, which are widely used for image classification and object detection. However, ResNet, U-Net, DeepLabv3, and MobileNet are mostly used for semantic segmentation (pixel) – based image (RGB) classification (; ). In agriculture, the high or low thermal dynamic changes during sunny–cloudy–rainy days and nights make it difficult to spatially process bulk thermal images, such as separation of leaf/canopy pixels from background pixels (; ). To solve this classification challenge, the author proposed a new DL architecture with several components [convolutions, grouped convolution, transposed convolution, batch normalization, rectified linear unit (ReLU), max pooling, depth concatenation, element-wise addition, 2D crop, softmax, and classification output layer]. The aim of this study was to develop a DL architecture and demonstrate the learning ability of the DL architecture to separate the leaf/leaf canopy from a greenhouse background (ground, windows, roof, etc.) in thermal images under various environmental conditions (sunny, cloudy, and rainy: day or/and night).

Materials and Methods

Thermal Image Acquisition System

The study was conducted in the greenhouse of the Vegetable and Flower Research Division, National Agriculture and Food Research Organization (NARO) in Tsukuba, Ibaraki, Japan. The Japanese cultivar “CF Momotaro York” (Takii Seeds Co., Ltd., Kyoto, Japan) of tomato (Solanum lycopersicum) grown in a Rockwool system was used for this experiment. The image data collection period ran from October 16, 2019 to September 30, 2020. The air temperature and relative humidity at 1.2 m above the ground surface ranged between 8.6 and 37.5°C, 32 and 96% from October 16, 2019 to April 16, 2020. The air temperature and relative humidity at 1.2 m above the ground surface ranged between 9.6 and 39.3°C, 34 and 95% from August 7, 2020 to October 28, 2020. Thermal images with 1040 × 780 pixel resolution (screen) were obtained, as shown in Figure 3, using a compact long-wave thermal camera [Thermo FLEX F50B-ONL (Nippon Avionics Co., Ltd., Yokohama, Japan)] under various environmental conditions at a minimum distance of 0.3 m from the top and maximum 2 m from the side of the targeted tomato plant.

Figure 3

All images were stored in a 24-bit thermal image format. The emissivity range of the thermal camera is 0.1 to 1. In this experiment, the emissivity of the tomato leaf was considered to be 0.98 (). The technical specifications of the thermal camera are listed in Table 1.

Table 1

Field of view, °FocusSpectral range, μmFrame rate, HzSensitivity, °CAccuracy
70°×70°Focus free8~147.50.05°C at 30°C±2°C or ±2% for 0–40°C (other conditions: ±4°C or ±4%)

Technical specification of the thermal camera (Thermo FLEX F50B-ONL).

Image Dataset Preparation

Figure 4 demonstrates the schematic diagram of the image dataset preparation for network analysis.

Figure 4

In total, 13,766 thermal images were obtained during this experiment. The thermal images were resized into their original spatial resolution (240 × 240 pixels), and denoising (manipulation of scale and emissivity) was performed by a thermal imaging processing software (InfReC Analyzer NS9500STD for F50, Nippon Avionics Co., Ltd.) to meet the network input dimension (240 × 240 pixels) requirements. Furthermore, Image Segmenter (Image Processing and Computer Vision Toolbox, MATLAB R2020a) was used to convert the pixels of each thermal image into two groups manually: leaf (255) and background (0) as shown in Figure 5A. These pixel values were stored in binary images. The frequency levels of the leaf and background pixels within the total thermal image datasets were 77 and 23%, respectively (Figure 5B). In this experiment, 60% of the randomly selected images (thermal images and binary images) were used for training, 20% for validation, and 20% for test purposes.

Figure 5

The image dataset was augmented to increase the amount and type of variation within the training image data to prevent overfitting and generalizing the model performance (Figures 6AE). Table 2 shows the number of image datasets used for deep learning analysis. First, we augmented the image data, including random reflection in the X and Y directions [(aug1)]. This dataset was used for the network performance study. Furthermore, for comparative analysis, we also augmented the thermal image dataset with the other four options (aug2), as shown in Table 3.

Figure 6

Table 2

ConditionOriginal datasetBinary datasetTrainingValidationTest
Total image number13,76613,7668,2602,7532,753

The number of image datasets used for the DL analysis.

Table 3

Augmentation optionsTraining optionVisualization
RandXreflection1Aug1Figure 6A
RandYreflection1
RandRotation[−90 90]Aug2Figure 6B
RandScale[1 1]
RandXScale[0.8 1.2]Figure 6C
RandYScale[0.8 1.2]
RandXShear[−20 20]Figure 6D
RandYShear[−20 20]
RandXTranslation[−10 20]Figure 6E
RandYTranslation[−10 20]

Properties of the augmented datasets for DL analysis.

Network Architecture

Figure 7 demonstrates the basic network architecture of the TheLNet270v1, which is a combination of the semantic segmentation-based network (convolution layers) and classification-based network (softmax). The convolution layer of the proposed network extracts the higher-level features from input images with multiple smaller filter sizes (3 × 3 × 3 × 32). The smaller filter size of the convolution layer has a strong generalization ability when the same types of objects within an image are conglutinated with each other (). This capability effectively improves network learning performance. According to , the ReLUs activation function added non-linearities to the model, converted values less than zero to zero for each element of the input, transformed the summed weighted input from the node into output, and allowed models to learn faster with higher accuracy. The batch normalization layer increases the network stability and normalizes the output of a previous activation layer by subtracting the batch mean and dividing by the batch SD (). introduced grouped convolution for training AlexNet with less powerful GPUs with limited RAM. It is also termed as convolutions in parallel as this layer separates input channels into groups by applying sliding convolution filters (vertically and horizontally), computing the input and weights, adding a bias, and finally combining the convolutions for each group independently (; ). We included grouped convolution to increase the width of the network without hampering computational power. According to and , the max-pooling layer simplifies the network complexity by compressing and extracting the main features, ensuring feature position and rotation invariance, and rotation reduced computing time. A 2D image cropping layer crops images at the center to explore contextual features (). The last convolution layer has two outputs corresponding to two classes with a ReLU activation followed by a batch normalization layer with 16 filters. The output of the last convolution layer is fed into the softmax layer for calculating the probability of the output classification layer. Finally, these expanded features are passed to the classification layer for classification (). Therefore, the depth of the DL architecture is fixed to 270 layers and accurately optimized based on training performance. The characteristics of the TheLNet270v1 architecture are shown in Table 4.

Figure 7

Table 4

Layers nameTotal number of layers
Image input1
Convolution31
ReLU83
Batch normalization71
Max pooling15
Transposed convolution49
Addition layer2
Grouped convolution4
Depth concatenation11
Crop2D1
Softmax1
Pixel-classification (output)1

Characteristics of the TheLNet270v1 architecture.

Network Parameters

The TheLNet270v1 was trained on a FUJITSU SHIHO Supercomputer equipped with TESLA V100-SXM2 32GB and CUDA version 10.2, DL, and parallel computing toolbox (MATLAB R2020a). The adaptive moment estimation (ADAM) algorithm was used to optimize the network weights. The transfer learning parameters applied for training the TheLNet270v1 were as follows: training option: Adam; validation frequency: 10; mini-batch size: 50/70/90/128/156/220/240/260/290/320; max epoch: 5/12/20/30/40; learn rate schedule: piecewise; shuffle: every-epoch; initial learn rate: 0.001; epsilon: 1e-08. ADAM was used to optimize the network weights. Table 5 shows the hyperparameter optimization parameter for the TheLNet270v1 training.

Table 5

Training options: adamExecution environment: parallel
Learn rate schedule: piecewiseValidation patience: Inf
Shuffle: every-epochEpsilon:1e-8
Verbose: falseInitial learn rate: 1.0000e-03
Validation frequency:10Learn rate drop factor: 0.1000
Gradient decay factor: 0.9000Learn rate drop period: 10
Squared gradient decay factor: 0.9990Gradient threshold method: l2norm
L2 regularization: 1.0000e-04Verbose frequency: 50
Gradient threshold: InfDispatch in background: 0
Sequence padding value: 0Reset input normalization: 1
Sequence length: longestSequence padding direction: right

Hyperparameter optimization parameter.

Comparative Analysis and Evaluation Metrics

Currently, MobileNetv2 is widely used in low-powered mobile devices for image recognition or classification tasks because of its simple network architecture and lower computational complexity (). first introduced ResNet with cross-layer connectivity in a CNN, which sped up the convergence of deep neural networks, solved the vanishing gradient problem by actively deploying special skip connections and a batch normalization layer and 20 and 8 times deeper than AlexNet and VGG. On the other hand, U-Net is mostly used in high-powered fixed devices because of its complex network architecture. It is widely used for biomedical image segmentation and classification purposes (). The bottleneck layer between the contracting and expanding paths of the U-Net architecture increased the network depth and was regularized by dropout to solve the overfitting issue during the network learning process (; ). stated that Deeplabv3plus employs atrous convolution or dilated convolutions in parallel or in cascade to extract dense features at multiple scales with better-stored information capability. TheLNet270v1 is designed so that it can be used in both low-powered mobile or high-powered fixed devices. There are several performance metrics such as training/validation/test accuracy (shows the percentage of correctly classified pixels), global accuracy (measuring ratio of correctly classified pixels to the total number of pixels), mean accuracy (measuring the percentage of correctly identified pixels for each class), confusion metrics, validation loss, training time, IoU/Jaccard index (measuring the amount of overlap per predicted class), weighted IoU (measuring the average IoU of each class), BF score (Boundary F1 – measuring the quality of the predicted boundary with the ground truth boundary), etc. are used for quantifying TheLNet270v1 accuracy and network efficiency. The same performance metrics were also evaluated on Deeplabv3plus (with a pretrained network MobileNetv2 and ResNet-50) and U-Net for comparative analysis.

Results and Discussion

Image datasets are augmented into two categories for network training. The augmented dataset1 and augmented dataset2, as shown in Table 6, are both used for performance study and comparative analysis.

Table 6

Image dataOriginal imageBinary imageTotal imageComments
Original dataset13,766 × 113,766 × 127,532-
Augmented dataset113,766 × 213,766 × 255,064Original condition + XYReflection
Augmented dataset213,766 × 613,766 × 682,596Original condition + XYReflection + Rotation + Scale + XYScale + XYShear + XY Translations

The image datasets for performance study and comparative analysis.

Feature Extraction and Activation for Visualization

Features extracted and visualized from the different depths of the TheLNet270v1 layers after completing the training are shown in Figure 8. Typical looking filters starting from the first layer in Figure 8B(I) show the colorful smooth pixels of each of the 64 filters, to noisy pixels in Figure 8B(II), and then slightly visible some features in Figure 8B(III). The last convolution layer in Figure 8B(IV) finally represents the visible pixel class. In Figures 8C(I,II), identical features of the grouped convolution layer in shallow depth are shown at different positions of an image. Figures 8DI reveal different structures of the feature maps within each filter and layer, and visualizations show that the feature map is activated on the foreground tomato leaf image, not the background objects. Finally, softmax (Figure 8J) gives a discrete probability for each class (leaf/leaf canopy and background), which is between 0 and 1, and the result is visualized in the pixel classification output layer (Figure 8K), where 1 (white color) means leaf/leaf canopy and 0 means background (black color).

Figure 8

Performance Metrics

Figure 9 shows the accuracy and loss of the training and validation datasets used to monitor the network overfitting issue. It is clearly visible that the model performs well on both training and validation data sets.

Figure 9

The pixel-level classification of thermal images by TheLNet270v1 was investigated. A validation accuracy of 95.22% was achieved with a minibatch size of 320, max epoch of 20, and training time of 94.15 min, shown in Figures 10A,B. Under the same conditions, the maximum IoU of 74 and 87% for leaf and background was achieved. During this time, a minimum validation loss of 12% was observed. The confusion matrix is given in terms of percentage and absolute number. It can be seen from the confusion chart in Figure 10C that the higher classification accuracies of 98.07, 98.06, and 98.07% for leaf and 85.89, 85.80, and 85.51% for the background achieved with the training, validation, and testing datasets (Table 6) and demonstrated that the network was well-trained.

Figure 10

Table 7 shows the test results of several other performance metrics such as global accuracy, mean accuracy, weighted IoU, and BF score. A higher value indicates better network performance.

Table 7

AccuracyGlobal accuracy, %Mean accuracy, %Mean IoU, %Weighted IoU, %Mean BFScore, %
Train metrics94.8591.9986.5090.3386.42
Validation metrics94.8291.6986.4390.2886.34
Test metrics94.8191.7586.2090.2686.36

The performance metrics for image datasets.

The classification accuracy of each class (leaf and background) is described in Table 8.

Table 8

AccuracyTrain metricsValidation metricsTest metrics
AccuracyIoUMean BFScoreAccuracyIoUMean BFScoreAccuracyIoUMean BFScore
Leaf97.2793.5790.9997.2493.5491.0397.2993.5691.10
Background86.7279.4281.7886.7079.3381.6086.2178.8481.56

The intersection over union (IoU) and BFScore for each class.

Figure 11A shows an example of a test image successfully segmented into two classes, in which the dark color area represents leaf and light color background. Figure 11B shows a tiny presence of false positives (magenta color). However, the boundary between leaf and background is marked as green color (true negatives), which described that further refinement is possible if we retrain the network with more image data or images with higher resolutions.

Figure 11

Comparative Metrics

It is evident from Table 9 that the TheLNet270v1 has a maximum depth layer of 270 with a lower total number of network parameters of 2e + 11, which is lower than Deeplabv3plus (ResNet50) and Deeplabv3plus (MobileNetv2). However, U-Net has a minimum of 46 layers with a higher total number of network parameters of 6e + 06 than TheLNet270v1. However, the training time for all networks (20 epoch, 220 minibatch sizes, and augmented dataset1) slightly differed.

Table 9

Network nameTotal layersTotal neuronsTotal weightsTotal biasesTotal parametersTraining time, min
Deeplabv3plus (ResNet50)2064.00E + 076.00E + 064.00E + 044.00E + 1280.51
Deeplabv3plus (MobileNetv2)1864.00E + 077.00E + 064.00E + 044.00E + 1294.29
U-Net468.00E + 077.00E + 063.00E + 036.00E + 0694.16
TheLNet270v12705.97E + 071.63E + 063.00E + 032.00E + 1194.15

Comparative statistics of the various network architectures.

Figure 12, Δ Performance (Eq. 1) demonstrated each evaluation metric’s positive and negative values with different image datasets. A negative value indicates an increase in the network performance, while a positive value is decreasing. The longer red arrow in the image indicates the volatile nature of the network due to the increase in the image dataset. From this, it is clear that Deeplabv3 (MobileNetv2) and TheLNet270v1 both show stable network performance despite increasing the number of images in the augmented dataset, as described in Table 6.

Figure 12

Test results vs. expected ground-truth (labeled) on the image-basis test dataset with IoU histogram are shown in Figures 13AD [Deeplabv3plus (ResNet50), Deeplabv3plus (MobileNetv2), U-Net, and TheLNet270v1], and the mean IoU of each class, as described in Table 10. The mean IoU of the leaf and background classes is indicated by the top bar in the image histogram. Figure 13 and Table 10 show that the difference in mean IoU is clearly noticeable for the network trained with augmented dataset1 and augmented dataset2. No noticeable changes are occurring for networks trained with different types of data sets. However, U-Net demonstrated IoU improvement with an increasing number of image datasets. The results revealed that leaves that counted the maximum number of pixels had lower IoU than the background with the least number of pixels. Further increasing the number of images within the same pattern or adding high-resolution images can improve the network performance ().

Figure 13

Table 10

Network nameAugmented data1Augmented data2Augmented data1Augmented data2Visualization
LeafBackground
Deeplabv3plus (ResNet50)0.740.740.870.85Figure 13A
Deeplabv3plus (MobileNetv2)0.720.720.860.85Figure 13B
U-Net0.440.660.520.82Figure 13C
TheLNet270v10.730.70.870.84Figure 13D

The mean IoU for comparative analysis.

Prediction Results

The prediction results of the independent image datasets are shown in Figure 14. Figure 14A represents the early morning with a sunny condition, Figure 14B represents the midday with a sunny–cloudy condition, and Figure 14C represents the midnight condition. These three sets of images were captured during September 2020 and were used to verify the network prediction efficiency. It is visualized that the TheLNet270v1 has better prediction ability compared with other networks.

Figure 14

Network Performance Verification With the IPPN Plant Phenotyping Image Dataset

As we described earlier, thermal images of the greenhouse-grown tomato plants were used for training the TheLNet270v1. This network successfully classified leaf/canopy and its background with higher accuracy, as shown in Figures 11, 14. We further investigated TheLNet270v1 performance using the IPPN plant phenotyping image dataset (leaf segmentation challenge component of the CVPPP workshop: CVPPP2017LSC-2017 and CVPPP2017LCC-2017; ), which included pot-cultivated Arabidopsis thaliana. First, we predicted the TheLNet270v1 output using image data from CVPPP2017LCC-2017 and CVPPP2017LSC-2017, as shown in Figures 15A,C. Subsequently, we trained the network with the CVPPP2017LSC-2017 image dataset (total images: 236, RGB) and then predicted again with the same image data from CVPPP2017LCC-2017 (Figures 15B,D). It is clearly visible that the TheLNet270v1 output, which is almost identical to the manually segmented binary image, is shown in Figure 15. Table 11 shows the TheLNet270v1 performance metrics.

Figure 15

Table 11

AccuracyMean IoU, %Weighted IoU, %Mean BFScore, %
Train metrics49.8699.7299.59
Validation metrics49.8499.6899.61
Test metrics49.9199.8199.81

The performance metrics for image datasets.

IoU, intersection over union.

Conclusion

This study introduced TheLNet270v1, a highly compact deep neural network (for mobile and non-mobile image classification) for classifying thermal images captured inside a greenhouse and demonstrating a higher classification accuracy. This paper also concludes a comparative analysis with other widely cited pre-trained networks for pixel-based classification, such as Deeplabv3plus (ResNet50), Deeplabv3plus (MobileNetv2), and U-Net, and found that TheLNet270v1 achieved a significantly better balance between accuracy and network efficiency. In our future work, we will apply the TheLNet270v1 network for on-site training, and output will be used for 24 h to monitor the relationships between plant growth and environmental conditions of the greenhouse. This network is suitable for the image with 240 × 240 pixels. However, to make it suitable for different pixel sizes, we consider modifying this network depending on the different image sizes in our future study.

Statements

Data availability statement

The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.

Author contributions

MI performed the architecture development and analysis and wrote the first draft of the manuscript. MI, NK, and UL conducted the field experiment. KT and NK verified the experimental results. All authors contributed to manuscript revision and approved the submitted version.

Funding

This work was supported by NARO under the “Environment optimization control system in plant factory facility using AI technology” Program (no. C11).

Acknowledgments

This research is the output of patented technology “A leaf temperature acquisition device, a crop growing system, a method for acquiring leaf temperature and a program for acquiring leaf temperature,” Japanese patent no. 2020-138804. The author gratefully acknowledges the technical support from Nippon Avionics Co., Ltd., Japan. We would like to acknowledge Tadahisa Higashide, NARO Institute of Vegetable and Flower Research.

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  • 1

    BengioY. (2009). Learning deep architectures for AI. Found Trends®. Mach. Learn. Available at: https://www.iro.umontreal.ca/~lisa/pointeurs/TR1312.pdf (Accessed October 14, 2020).

  • 2

    BhattaraiM.MartíNez-RamónM. (2020). A deep learning framework for detection of targets in thermal images to improve firefighting. IEEE Access.8, 8830888321. doi: 10.1109/ACCESS.2020.2993767

  • 3

    BlaschkeT. (2010). Object based image analysis for remote sensing. ISPRS J. Photogramm.65, 216. doi: 10.1016/j.isprsjprs.2009.06.004

  • 4

    BoulentJ.FoucherS.ThéauJ.St-charlesP. L. (2019). Convolutional neural networks for the automatic identification of plant diseases. Front. Plant Sci.10:941. doi: 10.3389/fpls.2019.00941

  • 5

    ChaerleL.Van Der StraetenD. (2000). Imaging techniques and the early detection of plant stress. Trends Plant Sci.5, 495501. doi: 10.1016/S1360-1385(00)01781-7

  • 6

    ChenC.MaY.RenG. A. (2019). Convolutional neural network with Fletcher–Reeves algorithm for hyperspectral image classification. Remote Sens.11, 1325. doi: 10.3390/rs11111325

  • 7

    ChoY.JulierS. J.MarquardtN.Bianchi-BerthouzeN. B. (2017). Robust tracking of respiratory rate in high-dynamic range scenes using mobile thermal imaging. Biomed. Opt. Express8, 44804503. doi: 10.1364/BOE.8.004480

  • 8

    ChoiY.KimN.HwangS.KweonI. S. (2016). “Thermal image enhancement using convolutional neural network,” in Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS’16; October 09–14, 2016; Daejeon.

  • 9

    ChristopherM.BelghithA.BowdC.ProudfootJ. A.GoldbaumM. H.WeinrebR. N.et al. (2018). Performance of deep learning architectures and transfer learning for detecting glaucomatous optic neuropathy in fundus photographs. Sci. Rep.8:16685. doi: 10.1038/s41598-018-35044-9

  • 10

    DauphinY. N.FanA.AuliM.GrangierA. (2017). “Language modeling with gated convolutional networks,” in Proceedings of the 34th International Conference on Machine Learning, PMLR’17; August 06–11, 2017; Sydney.

  • 11

    DongC.LoyC. C.HeK.TangX. (2016). Image super-resolution using deep convolutional networks. IEEE Trans. Pattern Anal. Mach. Intell.38, 295307. doi: 10.1109/TPAMI.2015.2439281

  • 12

    Fuentes-PachecoJ.Torres-OlivaresJ.Roman-RangelE.CervantesS.Juarez-LopezP.Hermosillo-ValadezJ.et al. (2019). Fig plant segmentation from aerial images using a deep convolutional encoder-decoder network. Remote Sens.11:1157. doi: 10.3390/rs11101157

  • 13

    GiustiA.CiresanD. C.MasciJ.GambardellaL. M.SchmidhuberJ. (2013). “Fast image scanning with deep max-pooling convolutional neural networks,” in Proceedings of the IEEE International Conference on Image processing, ICIP’13; September 05–18, 2013; Melbourne.

  • 14

    GrbovicZ.PanicM.MarkoO.BrdarS.CrnojevicV. (2019). “Wheatear detection in RGB and thermal images using deep neural networks,” in Proceedings of the International Conference on Machine Learning and Data Mining, MLDM’19; July 20–25, 2019; New York.

  • 15

    HeK.ZhangX.RenS.JianS. (2016b). “Identity mappings in deep residual networks,” in Proceedings of the European Conference on Computer Vision; October 11‒14, 2016. Amsterdam: ECCV.

  • 16

    HeK.ZhangX.RenS.SunJ. (2016a). “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR’16; June 27–30, 2016; Las Vegas.

  • 17

    HuangG.LiuZ.Van Der MaatenL.WeinbergerK. Q. (2017). “Densely connected convolutional networks,” in Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR’17; July 21–26, 2017; Honolulu, HI, United States.

  • 18

    IoffeS.SzegedyC. (2015). “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine learning; July 07–09, 2015. France: PMLR. Available at: http://proceedings.mlr.press/v37/ioffe15.pdf (Accessed October 19, 2020).

  • 19

    KhanA.SohailA.ZahooraU.QureshiA. S. (2020). A survey of the recent architectures of deep convolutional neural networks. Artif. Intell. Rev.53, 54555516. doi: 10.1007/s10462-020-09825-6

  • 20

    KraftM.WeigelH.MejerG.BrandesF. (1996). Reflectance measurements of leaves for detecting visible and non-visible ozone damage to crops. J. Plant Physiol.148, 148154. doi: 10.1016/S0176-1617(96)80307-5

  • 21

    KrizhevskyA.SutskeverI.HintonG. E. (2012). “Imagenet classification with deep convolutional neural networks,” in Proceedings of the Neural Information Processing Systems Conference; December 03–08, 2012. NV: NIPS.

  • 22

    KuenJ.KongX.WangG.TanY. P. (2018). “DelugeNets: Deep networks with efficient and flexible cross-layer information inflows,” in Proceedings of the IEEE International Conference on Computer Vision Workshop, ICCVW’17; October 22–29, 2017; Venice, Italy.

  • 23

    LarssonG.MaireM.ShakhnarovichG. (2016). “Fractalnet: Ultra-deep neural networks without residuals,” in Proceedings of the International Conference on Learning Rerepresentations, Toulon, ICLR’17; April 24–26, 2016; Toulon.

  • 24

    LiC.XiaW.YanY.LuoB.TangJ. (2019). Segmenting objects in day and night:edge-conditioned CNN for thermal image semantic segmentation. arXiv [Preprint]. Available at: https://arxiv.org/abs/1907.10303 (Accessed July 16, 2020).

  • 25

    LiliZ.DuchesneJ.NicolasH.RivoalR.BregerP. (1991). Détection infrarouge thermique des maladies du blé d’hiver (Infrared detection of winter-wheat diseases). EPPO Bull.21, 659672. doi: 10.1111/j.1365-2338.1991.tb01300.x

  • 26

    LiuJ.WangX. (2020). Early recognition of tomato gray leaf spot disease based on MobileNetv2-Yolov3 model. Plant Methods16:83. doi: 10.1186/s13007-020-00624-2

  • 27

    LópezA.Molina-AizF. D.ValeraD. L.PeñaA. (2012). Determining the emissivity of the leaves of nine horticultural crops by means of infrared thermography. Sci. Hortic.137, 4958. doi: 10.1016/j.scienta.2012.01.022

  • 28

    MaoX.ShenC.YangY.-B. (2016). “Image restoration using very deep convolutional encoder-decoder networks with symmetric skip connections,” in Proceedings of the Advances in Neural Information Processing Systems, NIPS’16 (Barcelona).

  • 29

    MinerviniM.FischbachA.ScharrH.TsaftarisS. A. (2015). Finely grained annotated datasets for image-based plant phenotyping. Pattern Recogn. Lett.81, 8089. doi: 10.1016/j.patrec.2015.10.013

  • 30

    NairV.HintonG. E. (2010). “Rectified linear units improve restricted Boltzmann machines,” in Proceedings of the 27th International Conference on Machine Learning, ICML’10; June 21–24, 2010; Haifa.

  • 31

    RazaS. E.PrinceG.ClarksonJ. P.RajpootN. M. (2015). Automatic detection of diseased tomato plants using thermal and stereo visible light images. PLoS One10:e0123262. doi: 10.1371/journal.pone.0123262

  • 32

    RonnebergerO.FischerP.BroxT. (2015). “U-net: convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015. Lecture Notes in Computer Science. eds. NavabN.HorneggerJ.WellsW.FrangiA. (Cham: Springer), 234241.

  • 33

    SaleemM. H.PotgieterJ.Mahmood ArifK. M. (2019). Plant disease detection and classification by deep learning. Plan. Theory8:468. doi: 10.3390/plants8110468

  • 34

    SalgadoeA. S. A.RobsonA. J.LambD. W.SchneiderD. (2019). A non-reference temperature histogram method for determining Tc from ground-based thermal imagery of orchard tree canopies. Remote Sens.11:714. doi: 10.3390/rs11060714

  • 35

    SchererD.MüllerA.BehnkeS. (2010). “Evaluation of pooling operations in convolutional architectures for object recognition,” in Artificial Neural Networks, ICANN 2010, Lecture Notes in Computer Science. eds. DiamantarasK.DuchW.IliadisL. S. (Berlin, Heidelberg: Springer), 92101.

  • 36

    ShinH.-C. C.RothH. R.GaoM.LuL.XuZ.NoguesI.et al. (2016). Deep convolution neural networks for computer aided detection: CNN architectures, dataset characteristics and transfer learning. IEEE Trans. Med. Imaging35, 12851298. doi: 10.1109/TMI.2016.2528162

  • 37

    SimonyanK.ZissermanA. (2014). “Two-stream convolutional networks for action recognition in videos,” in Proceedings of the 28th Conference on Neural Information Processing Systems, NIPS’14; December 08–13, 2014; Montréal, Canada.

  • 38

    SimonyanK.ZissermanA. (2015). “A very deep convolutional networks for large-scale image recognition,” in Proceedings of the 3rd International Conference on Learning Representations, ICLR’15; May 07–09, 2015; Sandiego.

  • 39

    SinghA. K.GanapathysubramanianB.SarkarS.SinghA. (2018). Deep learning for plant stress phenotyping: trends and future perspectives. Trends Plant Sci.23, 883898. doi: 10.1016/j.tplants.2018.07.004

  • 40

    SrivastavaR. K.GreffK.SchmidhuberJ. (2015). “Highway networks,” in Proceedings of the International Conference on Machine learning, ICML’15; July 06–11, 2015; Lille.

  • 41

    SzegedyC.LiuW.JiaY.SermanetP.ReedS.AnguelovD.et al. (2015). “Going Deeper with convolutions,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVP’15; June 06–12, 2015; Boston.

  • 42

    TongT.LiG.LiuX.GaoQ. (2017). “Image super-resolution using dense skip connections,” in Proceedings of the IEEE international conference on computer vision, ICCV’17; October 22–29, 2017 (Venice, Italy).

  • 43

    WongA.FamouriM.ShafieeM. J. (2020). AttendNets: tiny deep image recognition neural networks for the edge via visual attention condensers. arXiv [Preprint]. Available at: https://arxiv.org/abs/2009.14385v1 (Accessed July 16, 2020).

  • 44

    XavierG.BengioY. (2010). “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the 13th International Conference on Artificial Intelligence and Statistics; May 13–15, 2010; Sardinia: AISTATS.

  • 45

    ZhangX.HeL.MajeedY.WhitingM. D.KarkeeM.ZhangQ. (2018). A precision pruning strategy for improving efficiency of vibratory mechanical harvesting of apples. Trans. ASABE61, 15651576. doi: 10.13031/trans.12825

  • 46

    ZhangQ.LiuY.GongC.ChenY.YuH. (2020). Applications of deep learning for dense scenes analysis in agriculture: a review. Sensors20:1520. doi: 10.3390/s20051520

  • 47

    ZhangW.TangP.ZhaoL. (2019). Remote sensing image scene classification using CNN-CapsNet. Remote Sens.11:494. doi: 10.3390/rs11050494

Summary

Keywords

deep learning, network architecture, classification, segmentation, thermal image

Citation

Islam MP, Nakano Y, Lee U, Tokuda K and Kochi N (2021) TheLNet270v1 – A Novel Deep-Network Architecture for the Automatic Classification of Thermal Images for Greenhouse Plants. Front. Plant Sci. 12:630425. doi: 10.3389/fpls.2021.630425

Received

17 November 2020

Accepted

02 June 2021

Published

01 July 2021

Volume

12 - 2021

Edited by

Reza Ehsani, University of California, Merced, United States

Reviewed by

Pouria Sadeghi-Tehran, Rothamsted Research, United Kingdom; Yiannis Ampatzidis, University of Florida, United States

Updates

Copyright

*Correspondence: Md. Parvez Islam,

This article was submitted to Technical Advances in Plant Science, a section of the journal Frontiers in Plant Science

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics