ORIGINAL RESEARCH article

Front. Plant Sci., 16 December 2022

Sec. Sustainable and Intelligent Phytoprotection

Volume 13 - 2022 | https://doi.org/10.3389/fpls.2022.992789

Rodent hole detection in a typical steppe ecosystem using UAS and deep learning

  • 1. Agricultural Information Institute, Chinese Academy of Agricultural Sciences, Beijing, China

  • 2. Key Laboratory of Agricultural Blockchain Application, Ministry of Agriculture and Rural Affairs, Beijing, China

  • 3. State Key Laboratory for Biology of Plant Diseases and Insect Pests, Institute of Plant Protection, Chinese Academy of Agricultural Sciences, Beijing, China

  • 4. Institute of Grassland Research, Chinese Academy of Agricultural Sciences, Key Laboratory of Biohazard Monitoring and Green Prevention and Control in Artificial Grassland, Ministry of Agriculture and Rural Affairs, Hohhot, China

Abstract

Introduction:

Rodent outbreak is the main biological disaster in grassland ecosystems. Traditional rodent damage monitoring approaches mainly depend on costly field surveys, e.g., rodent trapping or hole counting. Integrating an unmanned aircraft system (UAS) image acquisition platform and deep learning (DL) provides a great opportunity to realize efficient large-scale rodent damage monitoring and early-stage diagnosis. As the major rodent species in Inner Mongolia, Brandt’s voles (BV) (Lasiopodomys brandtii) have markedly small holes, which are difficult to identify regarding various seasonal noises in this typical steppe ecosystem.

Methods:

In this study, we proposed a novel UAS-DL-based framework for BV hole detection in two representative seasons. We also established the first bi-seasonal UAS image datasets for rodent hole detection. Three two-stage (Faster R-CNN, R-FCN, and Cascade R-CNN) and three one-stage (SSD, RetinaNet, and YOLOv4) object detection DL models were investigated from three perspectives: accuracy, running speed, and generalizability.

Results:

Experimental results revealed that: 1) Faster R-CNN and YOLOv4 are the most accurate models; 2) SSD and YOLOv4 are the fastest; 3) Faster R-CNN and YOLOv4 have the most consistent performance across two different seasons.

Discussion:

The integration of UAS and DL techniques was demonstrated to utilize automatic, accurate, and efficient BV hole detection in a typical steppe ecosystem. The proposed method has a great potential for large-scale multi-seasonal rodent damage monitoring.

1 Introduction

Rodent infestation is one of the main biological hazards that seriously affect the health of grassland ecosystems (). In grassland ecosystems in Mongolian Plateau, Brandt’s vole (BV, Lasiopodomys brandtii) is the major pest, which is a small, seasonal breeding rodent species living in social groups and digging complex burrow systems with up to approximately 5,616 holes/ha in high-density areas (). Dense BV holes accelerated erosion and desertification in grasslands, resulting in mass herbage and forage loss in Inner Mongolia (). In addition, BV is also the intermediate host for many severe human infectious diseases (). Accurate and rapid detection of BV holes is an urgent need to evaluate the rodent population density for better ecosystem and human health protection.

Traditionally, grassland rodent hole detection mainly relied on field surveys. Field surveys can straightforwardly obtain rodent information but are time-consuming and labor-intensive (; ; ). Recently, UAS, which has a millimeter-level spatial resolution and can quickly collect multi-scale, multi-temporal images in real-time, has emerged as a promising alternative for rodent hole investigation. Nesting traces (burrows, mounds, tunnels, etc.) of many rodent species, i.e., Yellow Stepped Vole (Eolagurus luteus) (), Great Gerbil (Rhombomys optimus) (; ) and Plateau Pika (Ochotona curzoniae) (), were able to be identified through manual interpretation. Nevertheless, manual interpretation is still labor-intensive and time-consuming in regard to the large number of images generated by UAS. To solve this problem, scholars have tried to apply various man-machine interaction algorithms to count rodent holes in UAS imagery, such as maximum likelihood classification (), object-oriented classification (; ), support vector machine (SVM) (), and Sobel filter (). All these methods have limitations in effectiveness and efficiency when compared with deep convolutional neural networks.

Deep learning (DL) shows great potential in BV hole detection, benefiting from the application of automatic deep convolutional feature extraction (; ; ). Recent studies using DL and UAS images have been widely applied to small object detection, such as detecting birds () and mammals in the wild (; ; ), identifying weeds (), counting small plants (), and extracting vehicles (, ). So far, four previous studies have employed UAS images and DL in rodent hole detection. identified large gerbil holes (6-12 cm in diameter) in desert forests using You Only Look Once (YOLO)v3 and YOLOv3-tiny. detected holes of Plateau Pika (Ochotona curzoniae) with a diameter of 8-12 cm, which is a medium-sized rodent, using a Mask Region-based Convolutional Neural Network (R-CNN). successfully detected grassland rat holes (unspecified species) using R-CNN and improved Single Shot MultiBox Detector (SSD). detected Levant voles burrows (2.5-7.5 cm in diameter) in farmlands and found that YOLOv3 provided relatively accurate and robust results. Previous studies explored various algorithms to detect different rodent holes under various environments. However, no specific one has focused on small-sized rodent holes, e.g., BV holes (4-6 cm in diameter), in a complex typical steppe ecosystem.

The following issues are encountered in BV hole detection in a typical steppe ecosystem. First, the characteristics of BV make hole detection challenging. BV is small-sized, making them much more difficult to be detected from UAS images than other rodent holes. In addition, unlike other rodent species, e.g., Gerbillinae, BV digs holes in a different way that would not result in obvious excavated soil around holes. Visual features of other rodent holes cannot be utilized directly in BV holes. Furthermore, the typical steppe ecosystem, which is the main habitat of BV, has many factors that can impede detection. Animal droppings, hoofprints, and, most importantly, shadows and shades of grass and rocks make BV holes difficult to visually identify from images or even in the field. Different occlusion and illumination conditions at different times and seasons will lead to various spectral and geometric features of BV holes, thus requiring a robust detection method for different seasons, which refers to better generalizability. For example, more lush grass in summer will result in more occlusion in hole observations than in winter, thus making the detection more difficult. To sum up, BV hole detection in a typical steppe ecosystem requires a new dataset and a suitable detection method that can overcome the abovementioned issues in different seasons.

On account of the practical problems, this study aims to develop a specific UAS image dataset and a cost-effective and robust DL method in Brandt’s voles hole detection in a typical steppe ecosystem. We collected datasets in two different seasons: summer and winter, then investigated six DL-based object detection models, including three two-stage detectors and three single-stage detectors, to explore their accuracy, speed, and generalizability in BV hole detection.

2 Study area and data

Study areas (Figure 1) are located in East Uzhumuqin Banner (45°31′0″ N, 116°58′0″E) in Xilingol League, which is in the northeastern region of Inner Mongolia in China. Xilingol League is the main steppe habitat for Brandt’s vole (others are the Hulunbeir League of China, the Republic of Mongolia, and the Baikal Lake region of Russia) (). Rodent infestation occurs annually in Xilingol League and is associated with drought and ecological deterioration. In Xilingol Leaure, East Uzhumuqin Banner is the severely damaged region, with a total of 950 km2 area affected in 2021, where BV is the primary pest.

Figure 1

The experiment was implemented in a typical steppe in East Uzhumuqin Banner. As shown in Figures 2A, B, environmental conditions, especially grass conditions, are different between the two seasons. Images were collected in the same pasture in summer (September 9th-12th) and winter (November 1st-5th) in 2020. Both selected seasons have ecological significance. Brandt’s voles reproduce from March to August. The population of the species peaks in September (). In November, BV holes start clustering for the winter. Most holes become inactive and filled by soil, stones, grass, and snow, thus disappearing. The number of BV holes reaches the lowest in winter and performs as the population baseline for the next year (). Therefore, detection results in September can represent the magnitude of the BV disaster of the current year, and detection in November can help to determine the peak number of rodents in the next year. Images collected in September and November were used to generate two datasets (Dataset1 and Dataset2, respectively). Moreover, we combined these two datasets to generate Dataset3 as a comprehensive dataset. Detailed information about datasets is listed in Table 1.

Figure 2

Table 1

Dataset1Dataset2Dataset3
Collection dataNovember 1st to 5thSeptember 9th to 12thThe combination of Dataset1 and Dataset2
UAS devicesDJI Inspire 2 + DJI Zenmuse X5S professional gimbal camera+DJI 15mm Micro Four Thirds lensDJI Inspire 2 + DJI Zenmuse X5S professional gimbal camera+Olympus M.Zuiko 45mm/1.8 lens
Flight altitude8 meter15 meter
Patch size500*5001000*1000
Patch numbers258722184805
BV hole numbers341232796691

Datasets information.

Flight and data collection was conducted during the whole day from 8 am to 5 pm. The UAS utilized in this study to capture BV hole images was a DJI Inspire 2 equipped with a DJI Zenmuse X5S professional gimbal RGB camera. A DJI 15mm Micro Four Thirds lens and an Olympus M.Zuiko 45mm/1.8 lens were used to capture RGB images in summer and winter, respectively. In preliminary experiments, we explored different flight heights in the detection. We found that a BV hole can be recognized when its bounding box covers at least 30*30 pixels. Therefore, we chose 8m and 15m as the flight heights for the two UAS devices to ensure sufficient spatial resolution. Also, the vertical shooting angle was pre-determined for ortho rectification and mosaicking of images. Images captured by the two lenses have the same size (5280*3956 pixels).

3 Methods

The main method of this study is a supervised object detection approach (Figure 3). First, UAS images were collected in two different seasons and preprocessed. Two seasonal datasets and one combination dataset are generated, respectively. Next, to determine the optimal DL model in BV’s hole detection, we trained six representative object detection DL models, including three one-stage and three two-stage models. Finally, the performance of the models was assessed and compared. Detailed methods were presented in Sections 3.1 to 3.4.

Figure 3

3.1 Data preprocessing

In this section, raw images captured by UAS were preprocessed through selection, cropping, and labeling. First, images that contained BV holes were selected by experts to filter out invalid data. Then, full scenes of UAS images were cropped into fixed-size patches for further approach regarding computing memory limitation. Dataset1 and Dataset2 were manually cropped into patches with 500*500 pixels and 1000*1000 pixels, respectively. Overlaps between patches were avoided. Third, we labeled BV holes in the patches in LabelImg (), which is an open-source graphical image annotation tool, by delineating the bounding boxes of BV holes. Every BV hole sample in patches was selected and double-checked by experts. A total of 4805 images with 6691 BV holes were manually annotated for the training, validation, and test datasets. Before training, we implemented data augmentation to extend the datasets, including random scaling, random flipping, random cropping, and hue-saturation-value transformation.

3.2 Model training

For each dataset, patches were randomly split into three parts: 80% for training, 10% for validation, and 10% for testing. The training set was used to train the deep learning model; the validation set was used to validate the adopted improving tactics; the testing set was used to evaluate the performance of trained deep learning models.

3.2.1 Deep learning models

We tested six commonly used DL models, which can be grouped into two categories: two-stage detectors and one-stage detectors. Two-stage detectors, e.g., Faster R-CNN, Region-based Fully Convolutional Network (R-FCN), and Cascade R-CNN, conduct region proposal generation and object classification using two different networks. Alternatively, one-stage detectors, e.g., Single Shot MultiBox Detector (SSD), RetinaNet, and YOLOv4, treat object detection as a simple regression problem, thus running the above operations only using one network. Compared with two-stage detectors, one-stage models usually achieve lower detection accuracy but much faster speed. To determine an optimal method that can achieve a balance between accuracy and speed in detecting BV holes from the UAS images, we investigated three representative models from each category. The brief descriptions of models are presented below.

Three two-stage models are Faster R-CNN, R-FCN, and Cascade R-CNN. Faster R-CNN is a classical region-based deep detection model proposed in 2015 (). Faster R-CNN generates feature maps using the deep residual network ResNet-101, which contains 101 convolutional and pooling layers, proposed by . In the first stage, a Region Proposal Network (RPN) narrows the number of candidate object locations to a small number (e.g., 1~2k) by filtering out most background samples. In the second stage, the proposals from the first stage get features of equal size through Region of Interest (RoI) pooling and are sent to the classifier. After being classified into specific classes, the final object detection results will be provided with more accurate locations via bounding-box regression. This model employed a fully convolutional network, which simultaneously predicts object bounds and objectness scores at each position. It truly realized end-to-end training by introducing the basis of Fast R-CNN, which greatly improved the detection speed and accuracy ().

R-FCN is a fast approach in the two-stage approach category (; ). R-FCN also adopted ResNet-101 as the feature extractor. An RPN proposes candidate RoIs, which are then applied on the score maps, using a bank of specialized convolutional layers as the output. The use of position-sensitive score maps addressed the dilemma between invariance/variance on translation. All learnable layers are convolutional and are computed on the entire image. The architecture of R-FCN enables nearly cost-free region-wise computation and speeds up training and inference. It has achieved competitive results with a significantly faster detection speed than the Faster R-CNN.

Cascade R-CNN is a multi-stage object detection algorithm released at the end of 2017 (; ). ResNet-101 is also used as the feature extraction backbone in this model. Different from other models, in Cascade R-CNN, increasing thresholds of Intersection over Union (IoU), which is an indicator to judge the degree of overlap between predictions and labels, are trained in multiple cascaded detectors. The cascaded detectors were trained sequentially, where deeper stages are more sensitive against close false positives (). The inference speed after cascade may be slightly slower but within acceptable limits. Cascade R-CNN is conceptually straightforward, simple to implement, and can be combined, in a plug-and-play manner, with many detector architectures.

Three one-stage models are SSD, RetinaNet, and YOLOv4. SSD uses a set of predefined boxes of different aspect ratios and scales to predict the presence of an object in a certain image (; ). Particularly, it utilizes different target sizes to extract feature maps and encapsulates all computations in a single network. This design makes SSD easy to train and faster than two-stage models. VGG-16 was employed in this model for feature extraction.

RetinaNet is a one-stage detector that can achieve comparable accuracy to some two-stage models by using focal loss to solve the foreground-background class imbalance problem (). ResNet-101 is used in feature extraction. A Feature Pyramid Network (FPN) is proposed to construct a multi-scale feature pyramid from one single-resolution input image. RetinaNet is multi-scale, semantically strong at all scales, and fast to compute.

YOLOv4,the fourth version of YOLO, is a widely used, state-of-the-art, real-time object detection system (; ). YOLOv4 was proposed in 2020, which used novel CSPDarknet53 as a backbone and added universal algorithms, e.g., DropBlock Regularization. The Spatial Pyramid Pooling block was added over the CSPDarknet53 to increase the receptive field of the backbone features and separates the most significant context features. Instead of the FPN used in YOLOv3, PANet was used as the method of parameter aggregation from different backbone levels for different detector levels. Benefiting from the novel backbone and new features, YOLOv4 has enhanced learning capability and improved detection accuracy while assuring its positioning speed compared with YOLOv3. It also became easier to train on a single GPU ().

3.2.2 Model hyper-parameter settings

Model hyper-parameters, i.e., learning rate, batch size, iterations, and epochs, were adjusted during training. All models are trained with the Stochastic Gradient Descent (SGD) algorithm, and the optimal values of these hyper-parameters are listed in Table 2. At the end of the training, the validation loss reached a convergence state for all six models. The experiment was implemented on NVIDIA Tesla P100 GPU with an Inter(R) Xeon(R) Gold 6132 CPU with 16 G RAM. All the methods were implemented in PyTorch.

Table 2

ModelsFeature extraction networkBatch sizeLearning rateIterationsEpochs
Faster R-CNNResNet-101640.001500001282
R-FCN640.0012000005128
Cascade R-CNN640.0012000005128
RetinaNet160.00146875300
SSDVGG16320.0011200003096
YOLOv4CSPDarknet53160.00146875300

Model hyper-parameters settings.

3.3 Model evaluation

Nine indicators, including True Positive (TP), False Positive (FP), True Negative (TN), False Negative (FN) (Table 3), Recall, Precision, Average Precision (AP), F1-score, Average AP, Average F1-score, and Frames Per Second (FPS) were utilized to evaluate the performance of models.

Table 3

Actual classPredicted class
PositiveNegative
PositiveTPFN
NegativeFPTN

The confusion matrix for the possible outputs.

TP is an outcome where the model correctly predicts the positive class. Alternatively, TN is an outcome where the model correctly predicts the negative class. FP is an outcome where the model incorrectly predicts the positive class, and FN is an outcome where the model incorrectly predicts the negative class. True or false was determined by the threshold of intersection over union (IoU). IoU measures the overlap ratio between the detected object (marked by a bounding box) and the ground truth (an annotated bounding box). The threshold was set to 0.5, which means that a detection result is determined as true when IoU>=0.5.

We use Recall and Precision to evaluate the predictability of the BV hole detection model. Recall presents the ability to find all relevant instances in a dataset (Equation 1), and Precision presents the percentage of the instances which are correctly detected (Equation 2).

AP and F1-score were employed to comprehensively evaluate the results since Recall and Precision reflect only one aspect of the model’s performance. AP (Equation 3) is the area under the curve of Precision and Recall rate, which is an intuitive evaluation standard for the model accuracy and can be used to analyze the detection effect of a single category. F1-score (Equation 4) is the harmonic mean of precision and recall.

In addition, the Average AP and Average F1-score are the arithmetic mean of the APs and F1-scores among prediction results in all three datasets.

In addition, FPS (Equation 5) is used to assess the model efficiency. In FPS calculation, 150 patches act as the input of the trained model to obtain the T (total running time). FPS can be calculated by Equation 5 (it usually takes more time for the first picture to load the model, so the time of the first picture is not counted). Higher FPS indicates higher speed and better efficiency.

4 Results

4.1 Results from different models

Six DL object detection models were compared through the experimental results (Table 4). In terms of accuracy, the accuracies of the two-stage models were higher than those of one-stage models, except YOLOv4. Faster R-CNN achieved the highest accuracy with 0.905 in Average AP, seconded by YOLOv4 (0.872). RetinaNet had the lowest Average AP (0.681). Regarding the Average F1-score, Faster R-CNN also had the highest value, 0.86, followed by R-FCN (0.853) and YOLO v4 (0.852). SSD had the lowest Average F1-score, which is 0.662. As shown in Figure 4, Faster R-CNN and YOLOv4 achieved the best accuracies combining all datasets.

Table 4

CategoriesModelsAverage APAverage F1-scoreFPS
Two-stageFaster R-CNN0.9050.8614.62
R-FCN0.8270.8536.86
Cascade R-CNN0.8140.8452.81
One-stageSSD0.8000.66229.14
RetinaNet0.6810.7677.17
YOLOv40.8720.85210.62

Accurcies and speed of different models.

Figure 4

In terms of speed, the running speed of the two-stage models was significantly slower than that of the one-stage. SSD uses a shallow VGG-16 network as its backbone and had the fastest running speed, 29.14 frames/second. YOLOv4 using CSPDarknet53 follows, which had 10.62 frames/second. The FPS of other models used the ResNet-101 network as backbones were all below eight frames/second.

4.2 Results from seasonal datasets

We also compared the models’ performance among different datasets collected in two seasons (Table 5). Generally, all results from different models and datasets had acceptable accuracy, most of which had above 65% AP and F1-score. We performed a t-test for each pair of datasets. Detection results in early winter (Dataset1) were more accurate than those in summer (Dataset2), considering their AP and F1-score were significantly different at the 10% confidence level. In addition, based on the statistical test results, detection accuracy using Dataset3 was significantly lower than using Dataset1 and higher than using Dataset2. Among all models, Faster R-CNN and YOLOv4 were the two best models in terms of generalizability, regarding their low standard deviation of AP and F1-score among the three datasets (Figure 5).

Table 5

ModelsDatasetsTrue numberPredicted numberPrecisionRecallApF1-score
Faster
R-CNN
Dataset13544270.7960.9610.9450.871
Dataset23213460.8320.8970.8840.863
Dataset36467350.7990.9090.8870.850
R-FCNDataset13543700.8680.9070.8880.887
Dataset23213290.8210.8410.7860.831
Dataset36466650.8290.8530.8080.841
Cascade R-CNNDataset13543600.8830.8990.8980.891
Dataset23213170.8110.8070.7430.809
Dataset36466510.8330.8390.8010.836
SSDDataset13543050.8690.7490.8800.805
Dataset23211490.8320.3860.7530.527
Dataset36464190.8310.5390.7570.654
RetinaNetDataset13543540.8950.8950.8670.895
Dataset23211830.9010.5140.5120.655
Dataset36464820.8810.6560.6630.752
YOLOv4Dataset13543330.9130.8590.9050.885
Dataset23212850.8740.7760.8370.822
Dataset36465480.9250.7850.8740.849

The performances of different models using three datasets.

Figure 5

5 Discussion

In this paper, we developed the first Brandt’s voles hole detection method that can be utilized in a typical steppe ecosystem using UAS and DL. We established the first BV hole UAS image dataset, including samples in summer and winter. Detection results from six popular DL models were explored and compared. To our knowledge, this is the first UAS-based hole detection study for the specific species, Lasiopodomys brandtii. Advantages, findings, and limitations have been discussed below.

5.1 Advantages of UAS and DL models

Generally speaking, DL models based on UAS imagery had satisfactory results in BV hole detection. For example, using the model of Faster R-CNN and YOLOv4 to detect BV holes in UAS images, we can achieve a high Average AP, i.e., 0.905 and 0.872, which is a compelling output. More importantly, the proposed approach significantly improved the efficiency of the investigation. UAS-DL-based methods took less time and labor than traditional field survey methods. Specifically, taking the 0.25hm2 plot (the commonly used size for a manual survey plot) as an example, traditional manual methods require five or six people to spend about 1 hour. Repetitive counting in traditional methods may lead to a huge margin of error. In addition, human trampling during the investigation may cause destructive damage to grasslands. Therefore, it is not suitable for large-scale and periodic repeated monitoring. In monitoring by UAS, the aerial photography acquisition requires only one person and takes about 15 minutes for the same area (0.25hm2), which greatly improves the survey efficiency without damage. The running speed of SSD and YOLOv4 can achieve the FPS of 29.14 frames/second and 10.62 frames/second. For example, the proposed method using YOLOv4 needs only about 10 minutes for a 0.25hm2 plot to obtain the detection result. To this end, the proposed framework of UAS and DL models is an effective and efficient method for identifying BV holes.

5.2 Model comparison and selection

We compared the models from the following three perspectives: accuracy, running speed, and generalizability.

From the accuracy perspective, as shown in Table 4, Faster R-CNN and YOLOv4 were the two most accurate models with the highest average AP (0.904 and 0.872). It also should be noted that Faster R-CNN had the highest recall (0.961) while its precision was relatively low (0.796), which indicates that Faster R-CNN can detect more BV holes but may contain more false detections (Figure 6). YOLOv4 had the highest precision (0.913) and an acceptable recall (0.859).

Figure 6

From the running speed perspective, SSD and YOLOv4, which used VGG16 and CSPDarknet53 as their feature extraction networks, were the two fastest models with the highest FPS (29.14 and 10.62). All other models were used the ResNet-101 network as backbones, which has more parameters, thus required longer running time. However, SSD was excluded in practical applications since it had the lowest Average F1-score (0.662).

From the generalizability perspective, we focused on accuracy stableness, which refers to the variance of accuracy among different datasets. As mentioned in section 4.2, all models were performed better in winter than summer (Table 5). The main reason for this was that grass withers in winter, which leads to a much clearer view field. Less occlusion from grass in winter will lead to fewer missing BV holes in the detection. Additionally, fewer shadows and shades in winter will result in fewer FP.

Specifically, we discussed the advantages and disadvantages of each model one by one based their performance. Faster R-CNN had the highest accuracy considering AP and F1-score but relatively lower running speed. The RPN used to select candidate objects in the first stage of the model can utilize multi-scale feature information that improves performance in detecting small objects, e.g., BV holes. In contrast, much detailed information in RPN results in a long running time. In addition, it could detect the most BV holes among all models but may contain more false detections because fewer negative classes were sent to training after random sampling in RPN. R-FCN had the fastest running speed among two-stage models but a relatively lower Average AP, which indicates that features of BV holes are less sensitive to the problem that R-FCN majorly solved, i.e., the contradiction between translation variance and invariance. Cascade R-CNN had the lowest running speed thanks to its cascade architecture, with only 2.81 frames/second. However, its accuracy did not get satisfactory improvement after the cascade.

The efficiency of the one-stage models was significantly better than two-stage models. Benefiting from the VGG-16 network, SSD had the fastest running speed, which was more than ten times faster than Cascade R-CNN. However, the F1-score of SSD was the lowest, directly caused by the lowest recall due to the insufficient convolutional layers to extract features. Although RetinaNet had a higher F1-score than SSD benefiting from the focal loss function, its Average AP was the lowest among all models. The reason for this could be that RetinaNet pays more attention to difficult samples, thus leading to worse performance in easy and majority samples. YOLOv4 was ranked second both in accuracy and speed. Its one-stage architecture and the novel CSPDarknet53 feature extraction network reduced calculation when maintaining accuracy. In YOLOv4, the receptive field increases, and the size of the feature map decreases as the network deepens. Features and locations become abstract and fuzzy as well. While Faster R-CNN used FPN to construct a multi-scale feature pyramid for small objects, YOLOv4 had lower accuracy than Faster R-CNN in BV hole detection. On the other hand, YOLOv4 had a more balanced performance on missing BV hole detection and false detection than Faster R-CNN due to calculating the confidence loss for all positive and negative samples.

We further explored the generalizability of Faster R-CNN and YOLOv4 in a supplementary experiment. Specifically, we employed Dataset3, which is the most comprehensive dataset, as the training set, and tested the models on single-season datasets separately, i.e., Dataset1 and Dataset2. Experimental results showed that when using a more comprehensive training dataset in Faster R-CNN, the detection accuracy was significantly improved in both summer and winter (Table 6). Alternatively, when using YOLOv4, the detection accuracy was improved only in the winter dataset but decreased in the summer dataset. Furthermore, the accuracy improvement by Faster R-CNN was higher than by YOLOv4. To conclude, Faster R-CNN had a more accurate and consistent performance in two seasons.

Table 6

ModelsTraining datasetTestdatasetTrue numberPredicted numberPreci-sionRecallApF1-score
Faster R-CNNDataset3Dataset13544110.8670.9830.9770.921
Dataset3Dataset23213710.8410.9700.9510.901
YOLO-v4Dataset3Dataset13543280.9400.8700.9370.903
Dataset3Dataset23214200.6620.8670.7600.750

The performance in the generalizability test using different models.

In summary, when considering accuracy, Faster R-CNN and YOLOv4 were the two best models; when considering running speed, SSD and YOLOv4 were preferred; when considering generalizability, both Faster R-CNN and YOLOv4 were achieved acceptable accuracy in two seasons. Therefore, if taking into account accuracy, speed, and generalizability, the YOLOv4 was the best choice.

5.3 Uncertainty and limitations

There were several sources of uncertainty in the proposed UAS-DL-based BV hole detection method. First, noises like animal droppings, hoofprints, and shadows of grass and rocks were frequently appeared within the UAS images (Figure 7). Most of the falsely detected BV holes came from the misclassification of these noises (Figure 8). Second, various shapes and sizes of BV holes were brought uncertainty and error in the detection. In the complex wild environment, the shape and size of rodent holes depend on various factors, e.g.,the size of the rodents that live in, the rodents’ activity level, and the erosion degree. The lack of outlier-shaped samples lead to many FNs (Figure 9).

Figure 7

Figure 8

Figure 9

The limitations of this study are presented as follows. First, due to the top-view perspective, BV holes occluded by other objects cannot be detected from UAS. Second, small target detection still needs to be improved, especially in complex environments, e.g., typical steppe ecosystems. More advanced algorithms may improve the accuracy and consistency of BV hole detection. Besides the abovementioned future directions, real-time and large-scale BV detection studies are essential for better rodent monitoring. Moreover, spatial distribution and surrounding environment analysis are worthy of further exploration, which can provide more valuable advice on rodent disaster management and grassland protection.

6 Conclusions

A UAS-DL-based BV hole detection framework that can be used in different seasons was developed in this study. After comparing different DL models’ accuracy, speed, and generalizability in bi-seasonal datasets, we suggested an optimal model, YOLOv4, for BV hole detection in typical steppe ecosystems. In addition, we established a bi-seasonal BV hole UAS image dataset. To our knowledge, this is the first study that employs UAS images in BV hole detection. Furthermore, the seasonal effect was first considered and solved in rodent hole detection studies. The suggested model and dataset have a great potential for large-scale multi-temporal rodent hole detection and better management by the grassland ecological protection departments.

Statements

Data availability statement

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Author contributions

MD, DW and SL initiated the research idea and designed and conducted the experiments. MD finish the writing of this manuscript with the assistance of DW. DW and SL provided the financial and equipment support to make this study possible. CL and YZ provided important insights and suggestions on this study from the perspective of algorithms and plant protection experts. All authors contributed to the article and approved the submitted version.

Funding

This work was supported by the Inner Mongolia Science and Technology Project (Grant No. 2020GG0112, 2022YFSJ0010), Central Public-interest Scientific Institution Basal Research Fund of the Chinese Academy of Agricultural Sciences (Grant No. Y2021PT03), Central Public interest Scientific Institution Basal Research Fund of Institute of Plant Protection (S2021XM05), Science and Technology Innovation Project of Chinese Academy of Agricultural Sciences (CAAS-ASTIP-2016-AII).

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

  • 1

    AmatoG.CiampiL.FalchiF.GennaroC. (2019). Counting vehicles with deep learning in onboard uav imagery. 2019 IEEE Symposium Comput. Commun. (ISCC), 16. doi: 10.1109/ISCC47284.2019.8969620

  • 2

    BochkovskiyA.WangC. Y.LiaoH. (2020). Yolov4: Optimal speed and accuracy of object detection. arXivarXiv:2004.10934. doi: 10.48550/arXiv.2004.10934

  • 3

    BrownL. M.LacoJ. (2015). Rodent control and public health: A description of local rodent control programs. J. Environ. Health78, 2829.

  • 4

    CaiZ.VasconcelosN. (2018). Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 61546162. doi: 10.48550/arXiv.1712.00726

  • 5

    CuiB.ZhengJ.LiuZ.MaT.ShenJ.ZhaoX. (2020). YOLOv3 mouse hole recognition based on Remote sensing images from technology for unmanned aerial vehicle. Scientia Silvae Sinicae56 (10), 199208. doi: 10.11707/j.1001-7488.20201022

  • 6

    DaiJ.LiY.HeK.SunJ. (2016). R-fcn: Object detection via region-based fully convolutional networks. Adv. Neural Inf. Process. Syst.29.

  • 7

    EtienneA.AhmadA.AggarwalV.SaraswatD. (2021). Deep learning-based object detection system for identifying weeds using UAS imagery. Remote Sens.13 (24), 5182. doi: 10.3390/rs13245182

  • 8

    EzzyH.CharterM.BonfanteA.BrookA. (2021). How the small object detection via machine learning and UAS-based remote-sensing imagery can support the achievement of SDG2: A case study of vole burrows. Remote Sens.13 (16), 3191. doi: 10.3390/rs13163191

  • 9

    GirshickR. (2015). “Fast r-cnn,” in Proceedings of the IEEE international conference on computer vision. 14401448.

  • 10

    GuoX.YiS.QinY.et al. (2017). Habitat environment affects the distribution of plateau pikas: A study based on an unmanned aerial vehicle. Pratacultural Sci.34 (6), 13061313. doi: 10.11829/j.issn.1001-0629.2017-0090

  • 11

    HeydariM.MohamadzamaniD.ParashkouhiM. G.EbrahimiE.SoheiliA. (2020). An algorithm for detecting the location of rodent-made holes through aerial filming by drones. Arch. Pharm. Pract.1, 55.

  • 12

    HeK.ZhangX.RenS.SunJ. (2016). “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition. 770778.

  • 13

    HongS. J.HanY.KimS. Y.LeeA. Y.KimG. (2019). Application of deep-learning methods to bird detection using unmanned aerial vehicle imagery. Sensors19 (7), 1651. doi: 10.3390/s19071651

  • 14

    JintasuttisakT.LeonceA.Sher ShahM.KhafagaT.SimkinsG.EdirisingheE. (2022). “Deep learning based animal detection and tracking in drone video footage,” in Proceedings of the 8th International Conference on Computing and Artificial Intelligence. 425431.

  • 15

    KellenbergerB.MarcosD.TuiaD. (2018). Detecting mammals in UAV images: Best practices to address a substantially imbalanced dataset with deep learning. Remote Sens. Environ.216, 139153. doi: 10.1016/j.rse.2018.06.028

  • 16

    LeCunY.BengioY.HintonG. (2015). Deep learning. Nature.521 (7553), 436444. doi: 10.1038/nature14539

  • 17

    LiG.HouX.WanX.ZhangZ. (2016a). Sheep grazing causes shift in sex ratio and cohort structure of brandt's vole: Implication of their adaptation to food shortage. Integr. Zoology11, 7684. doi: 10.1111/1749-4877.12163

  • 18

    LinT. Y.GoyalP.GirshickR.HeK.DollárP. (2017). Focal loss for dense object detection. IEEE Trans. Pattern Anal. Mach. Intell.PP (99), 29993007. doi: 10.1109/ICCV.2017.324

  • 19

    LinT. Y.MaireM.BelongieS.HaysJ.PeronaP.RamananD.et al. (2014). Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740755). (Springer, Cham).

  • 20

    LiuX. H. (2022). Discrepancy, paradox, challenges, and strategies in face of national needs for rodent management in China. J. Plant Prot.49 (01), 407414. doi: 10.1007/978-3-319-10602-1_48

  • 21

    LiuW.AnguelovD.ErhanD.SzegedyC.ReedS.FuC. Y.et al. (2016). “Ssd: Single shot multibox detector,” in European Conference on computer vision (Cham: Springer), 2137. doi: 10.1007/978-3-319-46448-0_2

  • 22

    LiuL.OuyangW.WangX.FieguthP.ChenJ.LiuX.et al. (2020). Deep learning for generic object detection: A survey. Int. J. Comput. Vision128 (2), 261318. doi: 10.1007/s11263-019-01247-4

  • 23

    LiG.YinB.WanX.WeiW.WangG.KrebsC. J.et al. (2016b). Successive sheep grazing reduces population density of Brandt’s voles in steppe grassland by altering food resources: a large manipulative experiment. Oecologia180(1), 149159. doi: 10.1007/s00442-015-3455-7

  • 24

    MaT.ZhengJ.WenA.ChenM.MuC. (2018a). Group coverage of burrow entrances and distribution characteristics of desert forest-dwelling Rhombomys opimus based on unmanned aerial vehicle (UAV) low-altitude remote sensing: A case study at the southern margin of the Gurbantunggut Desert in Xinjiang. Acta Ecologica Sin.38 (3), 953963. doi: 10.11707/j.1001-7488.20181021

  • 25

    MaT.ZhengJ.WenA.ChenM.LiuZ.. (2018b). Relationship between the distribution of rhombomys opimus holes and the topography in desert forests based on low-altitude remote sensing with the unmanned aerial vehicle (UAV) : A case study at the southern margin of the gurbantunggut desert in Xinjiang, China. Scientia Silvae Sinicae54 (10), 180188. doi: 10.5846/stxb201612142571

  • 26

    MountrakisG.LiJ.LuX.HellwichO. (2018). Deep learning for remotely sensed data. ISPRS J. Photogramm. Remote Sens.145, 12. doi: 10.1016/j.isprsjprs.2018.08.011

  • 27

    OhS.ChangA.AshapureA.JungJ.DubeN.MaedaM.et al. (2020). Plant counting of cotton from UAS imagery using deep learning-based object detection framework. Remote Sens.12 (18), 2981. doi: 10.3390/rs12182981

  • 28

    PengJ.WangD.LiaoX.ShaoQ.SunZ.YueH.et al. (2020). Wild animal survey using UAS imagery and deep learning: modified faster r-CNN for kiang detection in Tibetan plateau. ISPRS J. Photogrammetry Remote Sens.169, 364376. doi: 10.1016/j.isprsjprs.2020.08.026

  • 29

    RazaviS.DambandkhamenehF.AndroutsosD.DoneS.KhademiA. (2021). Cascade R-CNN for MIDOG Challenge. In International Conference on Medical Image Computing and Computer-Assisted Intervention (pp. 8185). Springer, Cham. doi: 10.1007/978-3-030-97281-3_13

  • 30

    RedmonJ.DivvalaS.GirshickR.FarhadiA. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 779788.

  • 31

    RenS.HeK.GirshickR.SunJ. (2015). Faster r-cnn: Towards real-time object detection with region proposal networks. Adv. Neural Inf. Process. Syst.28.

  • 32

    ShiD. (2011). Studies on selecting habitats of brandt's voles in various seasons during a population low[J]. Acta Theriologica Sin.6 (4), 287. doi: 10.16829/j.slxb

  • 33

    SovianyP.IonescuR. T. (2018). Optimizing the trade-off between single-stage and two-stage deep object detectors using image difficulty prediction. In 2018 20th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC) pp. 209214. (IEEE). doi: 10.1109/SYNASC.2018.00041

  • 34

    SunD.NiY.ChenJ.AbuduwaliZhengJ. (2019). Application of UAV low-altitude image on rathole monitoring of eolagurus luteus. China Plant Prot.39 (4), 3543. doi: 10.3969/j.issn.1672-6820.2019.04.006

  • 35

    TsangS. H. (2019) Review: R-FCN–Positive-Sensitive score maps (Object detection). Available at: https://towardsdatascience.com/review-r-fcn-positive-sensitivescore-maps-object-detection-91cd2389345c.

  • 36

    TzutalinD. (2016) LabelImg is a graphical image annotation tool and label object bounding boxes in images. Available at: https://github.com/tzutalin/labelImg.

  • 37

    WangD.LiN.TianL.RenF.LiZ.ChenY.et al. (2019). Dynamic expressions of hypothalamic genes regulate seasonal breeding in a natural rodent population. Mol. Ecol.28 (15), 35083522. doi: 10.1111/mec.15161

  • 38

    WanJ.JianD.YuD. (2021). Research on the method of grass mouse hole target detection based on deep learning. J. Physics: Conf. Ser.Vol. 1952, No. 2, 022061. doi: 10.1088/1742-6596/1952/2/022061

  • 39

    WanX.LiuW.WangG.ZhongW. (2006). Seasonal changes of the activity patterns of brandt’s vole (Lasiopodomys brandtii) in the typical steppe in inner Mongolia. Acta Theriologica Sin.26 (3), 226. doi: 10.3969/j.issn.1000-1050.2006.03.003

  • 40

    WenA.ZhengJ.ChenM.MuC.MaT. (2018). Monitoring mouse-hole density by rhombomys opimus in desert forests with UAV remote sensing technology. Scientia Silvae Sinicae54 (4), 186192. doi: 10.11707/j.1001-7488.20180421

  • 41

    XuanJ.ZhengJ.NingY.MuC. (2015). Remote sensing monitoring of rodent infestation in grassland based on dynamic delta wing platform (in Chinese). China Plant Prot. (2), 4. doi: 10.3969/j.issn.1672-6820.2015.02.014

  • 42

    XuY.YuG.WuX.WangY.MaY. (2016). An enhanced viola-Jones vehicle detection method from unmanned aerial vehicles imagery. IEEE Trans. Intelligent Transportation Syst.18 (7), 18451856. doi: 10.1109/TITS.2016.2617202

  • 43

    ZhangZ. B.WangZ. (1998). Ecology and management of rodent pests in agriculture. (Beijing: Ocean Press)

  • 44

    ZhaoX.YangG.YangH.XuB.WangY. (2016). Digital detection of rat holes in inner Mongolia prairie based on UAV remote sensing data. In Proceedings of the 4th China Grass Industry Conference.

  • 45

    ZhongW.WangM.WanX. (1999). Ecological management of brandt’s vole (Microtus brandti) in inner Mongolia, China. Ecologically-based Rodent Management. ACIAR Monograph59, 119214.

  • 46

    ZhongW.WangG.ZhouQ.WangG. (2007). Communal food caches and social groups of brandt's voles in the typical steppes of inner Mongolia, China. J. Arid Environments68, 398407. doi: 10.1016/j.jaridenv.2006.06.008

  • 47

    ZhouX.AnR.ChenY.AiZ.HuangL. (2018). Identification of rat holes in the typical area of“Three-river headwaters”region by UAV remote sensing. J. Subtropical Resour. Environ.13 (4), 8592. doi: 10.3969/j.issn.1673-7105.2018.04.013

  • 48

    ZhouS.HanL.YangS.WangY.GenxiaY.NiuP.et al. (2021). A study of rodent monitoring in ruoergai grassland based on convolutional neural network. J. Grassland Forage Sci. (02), 1525. doi: 10.3969/j.issn.2096-3971.2021.02.003

  • 49

    ZouX. (2019). “A review of object detection techniques,” in 2019 International Conference on Smart Grid and Electrical Automation (ICSGEA).

Summary

Keywords

rodent monitoring, mouse hole detection, grassland protection, unmanned aircraft vehicle (UAV), object detection

Citation

Du M, Wang D, Liu S, Lv C and Zhu Y (2022) Rodent hole detection in a typical steppe ecosystem using UAS and deep learning. Front. Plant Sci. 13:992789. doi: 10.3389/fpls.2022.992789

Received

13 July 2022

Accepted

29 November 2022

Published

16 December 2022

Volume

13 - 2022

Edited by

Ozgur Batuman, University of Florida, United States

Reviewed by

Yongxin Liu, Embry–Riddle Aeronautical University, United States; Karansher Singh Sandhu, Bayer Crop Science (United States), United States

Updates

Copyright

*Correspondence: Dawei Wang, ; Shengping Liu,

This article was submitted to Sustainable and Intelligent Phytoprotection, a section of the journal Frontiers in Plant Science

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics