ORIGINAL RESEARCH article

Front. Environ. Sci., 13 July 2026

Sec. Environmental Informatics and Remote Sensing

Volume 14 - 2026 | https://doi.org/10.3389/fenvs.2026.1707781

Automatic acquisition method of ground vegetation conditions based on U-Net and clustering algorithm in remote sensing monitoring

  • Alpha Earth Observation, Glasgow, United Kingdom

Abstract

The current agricultural and forestry remote sensing monitoring relies on low-level features to obtain information about the specific conditions of ground vegetation, and the quality of the obtained information is not high. In response to this situation, the study proposes an automatic acquisition method for ground vegetation conditions based on the U-shaped Network (U-Net) and clustering algorithm. The study first employs an improved U-Net model to classify and predict remote sensing images, distinguishing between vegetation areas and non-vegetation areas. Subsequently, the simple linear iterative clustering algorithm is utilized to perform fine-grained segmentation on the remote sensing images. Finally, the classification prediction results are combined with the fine-grained segmentation results to generate label values for each pixel within the segmentation area. This process yields a more precise vegetation distribution map. The outcomes revealed that the classification accuracy of the improved U-Net reached 95.46%, and the fine-grained segmentation accuracy was improved to 89.32%. In practical applications, the information extraction efficiency of this method was relatively fast, reaching 25.4 frames per second. Furthermore, the intersection over union and F1 score of the extracted results were 0.89 and 0.93, respectively. This indicates that while improving the quality of vegetation information acquisition, this method significantly enhances the processing speed, providing efficient and accurate technical support for agricultural and forestry vegetation management.

1 Introduction

The rapid advancement of remote sensing (RS) technology has presented unprecedented opportunities for large-scale and efficient monitoring of ground vegetation conditions. Particularly in fields such as agricultural and forestry resource management, ecological environment protection, and climate change research, accurately acquiring information on vegetation coverage, health status, and distribution patterns is of utmost importance (; ). High-resolution RS image (RSI), such as data from the Gaofen No.2 (GF-2) satellite, can provide rich surface textures and spectral details, laying a data foundation for refined vegetation identification (). However, in current agricultural and forestry RS monitoring practices, obtaining information on the specific conditions of ground vegetation still faces challenges. Many traditional methods primarily rely on pixel spectral information or preset low-level features (such as normalized difference vegetation index, texture features, etc.) for classification (). These methods often struggle to distinguish vegetation of different ages, health conditions, or site conditions in areas with complex terrain features and similar spectral characteristics. This can result in low-quality information, blurred boundaries, frequent misclassifications, and omissions. Ultimately, these methods fail to meet the management needs of modern precision agriculture and forestry ().

In recent years, both domestic and international scholars have made significant achievements in vegetation extraction, phenological monitoring, and related ecological impact studies. These studies provide valuable references for research methods and ideas. Minyu Wang et al. systematically summarized the extraction, validation, and product development of RS vegetation phenological parameters. The study examined the use of traditional vegetation indices for phenological monitoring. It analyzed relevant preprocessing procedures and extraction algorithms and reviewed multi-source validation methods and the current status of product release. In the future, the study noted, it will be important to significantly improve algorithm stability and data quality, lessen the reliance on human experience for models, and increase the usefulness of vegetation products (). Marra E et al. started with the impact of wood extraction operations on soil characteristics and vegetation growth environment. Through tractor uphill and downhill operation experiments combined with soil compaction and rutting measurement techniques, the study found that downhill hauling had an advantage in reducing soil disturbance. Moreover, the pressure was mainly concentrated in the wheel track area, providing empirical support for understanding the basic environment for vegetation growth (). Turyahabwe R et al. focused on the destructive impact of human activities on vegetation ecology in the absence of RS. Their study systematically evaluated the impact of traditional brick making on wetland vegetation degradation, water quality deterioration, and deforestation in the Goma region through a combination of questionnaire surveys, interviews, and field observations. It also pointed out that while bringing certain social and economic benefits, supervision, and policy support should be strengthened to ensure ecological sustainability ().

With the rapid development of deep learning (DL) methods in the field of RS, related research has gradually become a new trend in vegetation extraction and ecological monitoring. Hamedianfar A et al. reviewed the current applications and challenges of DL in forestry and vegetation extraction. They emphasized DL’s potential for multi-scale feature learning and pointed out that data quality, model architecture selection, and temporal feature fusion are urgent issues that need to be addressed (). Ottoni A L C et al. investigated the application of deep learning for vegetation image recognition, specifically utilizing the VGG-16 architecture and transfer learning strategies. Through a comprehensive hyperparameter tuning case study, the influence of various configurations was analyzed. The findings indicated that optimizing these parameters significantly enhances the model’s performance in accurately classifying vegetation targets (). Lou X et al. conducted additional research on crown extraction for unmanned aerial vehicle (UAV) imagery. They used three target detection algorithms to automatically identify the widths of pine tree crowns. This method provided an efficient alternative to traditional ground surveys ().

For the assessment of vegetation growth status, Wang Y P et al. proposed using UAV multispectral imagery combined with the nitrogen index (N-Index) to estimate the nitrogen content in rice during its vegetative growth phase. By selecting optimal vegetation indices (such as the normalized difference red edge index (NDRI) and the red edge chlorophyll index (RECI)) and growth stages, they revealed the spatial heterogeneity of nitrogen status and provided data support for precision fertilization (). Ganju, N. K. et al. quantified the proportion of salt marsh vegetation using a Landsat-based method that compares unvegetated and vegetated swamps. They achieved monitoring of intertidal wetland area dynamics and vulnerability zoning through multi-scale calibration and validation. This provided a reference for wetland ecological restoration (). Hakam O et al. utilized RS and meteorological indices to assess agricultural drought and crop health changes in the Lower Sidi Bouregreg Basin of Morocco between 1984 and 2016. The research findings revealed a significant deterioration in crop health at the beginning of the 21st century. Additionally, the results exhibited a stronger correlation between yield and the temperature condition index and the standardized precipitation evapotranspiration index. This implies that yield is more susceptible to variations in temperature ().

Although deep learning has demonstrated significant potential in vegetation monitoring, current methodologies remain constrained by insufficient multi-scale feature representation, low accuracy in fine-grained segmentation of complex terrains, and limited generalization across different scenarios. To bridge these gaps, this study proposes a deep clustering model for automatic vegetation status recognition (DC-AVSR). This approach improves the U-Net architecture by integrating the atrous spatial pyramid pooling (ASPP) module to enhance feature perception at various scales. Furthermore, a Markov distance-based binary K-means clustering algorithm is introduced to precisely distinguish internal vegetation differences, such as health status and forest age gradients. Consequently, this integrated method offers an efficient and accurate automated solution for the refined management and dynamic assessment of agricultural and forestry vegetation.

2 Automatic acquisition method of ground vegetation conditions based on U-Net and clustering algorithm

2.1 Construction of a RSI classification and prediction model based on U-Net

The universality and accuracy of the method for precise and efficient RS monitoring of ground vegetation largely rely on high-quality data sources and standardized data preprocessing processes (; ). The study selects the Hebei Saihanba mechanized forest farm field as the core research area. The satellite cloud image of the Hebei Saihanba mechanized forest farm field is displayed in Figure 1.

FIGURE 1

In Figure 1, the internal topography of the forest farm is slightly undulating, mainly consisting of plateaus, mountains, and hills. Although the main vegetation type is relatively single, due to staged and batch afforestation, the forest farm presents a patchy mosaic distribution pattern with different forest ages (young forest, middle-aged forest, mature forest) and different site conditions (sunny slope, shady slope, different altitude). This pattern results in significant differences in tree height, crown width, density, and chlorophyll content between larch and Pinus sylvestris. Manifested as subtle changes in spectral and texture features in RSIs. The study selects PMS (panchromatic/multispectral) sensor images from the domestically produced GF-2 satellite as the main data source. The key technical parameters of the selected GF-2 imaging are displayed in Table 1.

TABLE 1

ParameterIndicator
Orbital altitude631 km (Sun-synchronous, recurrent orbit)
Descending node local time10:30 a.m.
Revisit period5 days (±35° off-nadir)
Spatial resolutionPanchromatic (PAN): 0.81 m; multispectral (MS): 3.24 m
Spectral bandsB1 (Blue): 0.45–0.52 μm
B2 (Green): 0.52–0.59 μm
B3 (Red): 0.63–0.69 μm
B4 (Near-Infrared, NIR): 0.77–0.89 μm
Swath width45 km
Data levelLevel-1 (systematic radiometric and geometric correction applied)

Key technical parameters of GF-2 imagery.

The dataset consists of multi-date L1 level image products with cloud cover below 5%, acquired continuously during the vegetation growth seasons (June–September) of 2023 and 2024. These multi-temporal images were compiled to ensure high-quality, cloud-free coverage and to capture representative vegetation characteristics across different phenological stages within the core area of the Saihanba forest farm. Radiometric calibration is the process of converting the dimensionless raw grayscale values (GSVs) (digital number, DN) recorded by sensors into radiance values with practical physical significance. This is the basis for subsequent quantitative analysis such as atmospheric correction and vegetation index calculation (). The conversion formula is usually a linear gain offset model, as shown in Equation 1.

In Equation 1, represents the spectral radiance at the entrance pupil of the band satellite sensor. is the original GSV of the corresponding band in the image. is the scaling gain (slope). is the calibration offset. The two key parameters and can be obtained from the metadata file that comes with the imaging product. The study utilizes the “Radiometric Calibration” tool of ENVI software to automatically read metadata information and complete radiometric calibration for all multispectral bands. To achieve high-precision recognition of vegetation areas in complex backgrounds, a U-Net deep CNN is used as the core segmentation model. Clustering algorithms are then employed to further divide the internal state of vegetation (). The architecture of the U-Net model consists of a symmetric encoding path for multi-scale feature extraction and a decoding path for spatial resolution recovery, integrated via skip connections to maintain fine-grained edge details, as illustrated in Figure 2.

FIGURE 2

In each layer of convolution operation (CO), the input FM is displayed as , the convolution kernel (CK) weight is , the bias is , the convolution output is , and the output after activation function processing is . The expression for this process is shown in Equation 2.

In Equation 2, represents the standard CO. The size of the CK in the network is often set to , and the FM size is kept consistent through zero padding. The activation function uses ReLU, which is defined as Equation 3.

Equation 3 can effectively alleviate gradient vanishing and improve the convergence speed of the model. At the encoding end, after two convolutions, perform an Max Pooling operation to halve the spatial resolution and double the number of feature channels to obtain deeper abstract features. The decoding path gradually restores the spatial size of the FM through transpose convolution (). Let the input feature of the -th layer in the decoding stage be , its transpose CO can be expressed as shown in Equation 4.

In Equation 4, represents the transpose CO, and is the decoding CK parameter. The FM output by transpose convolution is concatenated with the FM of the corresponding encoding layer. This skip connection enables the model to fully utilize low-level edge texture information while maintaining global semantics, thereby significantly improving the segmentation accuracy of vegetation boundaries. Finally, the output layer compresses the number of feature channels to 1 through convolution. The Sigmoid function is used to map the results to probability values between 0 and 1, as shown in Equation 5.

In Equation 5, is the probability that pixel belongs to the vegetation category. is the feature value of the pixel in the output layer. When , it is determined that the pixel is vegetation, otherwise it is non-vegetation. A weighted combination of the cross entropy loss function and dice loss is chosen as the optimization goal in an attempt to address the problem of unequal proportions between the vegetation and non-vegetation categories (). Cross entropy loss characterizes the difference between pixel level probability prediction and true labels, as shown in Equation 6.

In Equation 6, is the total quantity of pixels. is the predicted probability of the pixel. corresponds to the real label (vegetation is 1, non-vegetation is 0). Equation 7 illustrates how dice loss directly quantifies the extent of overlap between the true and anticipated areas.

Combining Equations 6, 7, the final loss function proposed in the study is shown in Equation 8.

In Equation 8, is the weight hyperparameter. In the study, a value of 0.5 is taken to balance the requirements of region overlap and pixel classification accuracy.

2.2 FGS of RSIs based on clustering algorithm

After completing the construction of a RSI classification and prediction model based on U-Net, further research is conducted on FGS of the segmentation results to identify health differences, forest age gradients/potential degradation phenomena within vegetation areas. Due to the insufficient ability of distinguishing vegetation from non-vegetation to reflect the internal structure and state of land cover, research has introduced clustering algorithms based on segmentation results to achieve higher resolution state partitioning of vegetation areas (). Therefore, a FGS process based on RSIs is constructed. This process starts from multi-source high-resolution images and performs multidimensional feature extraction and encoding on the segmented vegetation areas. Further automated state grouping is achieved through an improved binary K-means clustering algorithm. The feature extraction stage comprehensively utilizes the original multispectral image and segmentation mask information to obtain feature vectors covering multiple aspects such as spectrum, texture, and structure. Among them, the spectral features are centered around the normalized vegetation index and calculated using Equation 9. Moreover, it provides multi-dimensional input for clustering by combining texture indicators and structural information ().

According to Equation 9, displays the reflectance in the near-infrared band. is the reflectance of the red light band. To compute indicators like energy, contrast, homogeneity, and others that indicate the intricacy of canopy structure, texture characteristics are based on gray level co-occurrence matrices (). The structural features are extracted through local binary patterns and edge gradient operators to extract forest texture and edge information. These features are normalized and input into a feature encoding network for quantization and high-dimensional embedding representation. Figure 3 shows the schematic process of feature quantization and prototype matching network.

FIGURE 3

In Figure 3, after the multi-source input features undergo quantization processing, local textures and spectral patterns are expressed more deeply through entity encoders. The prototype matching network maps different vegetation samples to prototype centers in high-dimensional feature space, enhancing intra-class consistency and improving inter-class separability. After feature extraction and encoding are completed, a system architecture for vegetation condition clustering analysis is developed to integrate the entire process of data storage, preprocessing, feature organization, and clustering calculation. The schematic diagram of the system architecture for vegetation feature extraction and clustering analysis is shown in Figure 4.

FIGURE 4

In Figure 4, the underlying data persistence module is responsible for storing multi-source RSIs and extracted feature tables, including reader information tables, resource usage tables, etc., abstracted as feature data tables in vegetation monitoring tasks (such as spectral feature tables, texture feature tables, and vegetation health index tables). Transform raw data into structured features through ETL process and pass it to the data preprocessing layer. The preprocessing layer denoises, standardizes, and selects features to provide a unified input format for subsequent clustering algorithms. The core of the system is the multi view binary K-means clustering algorithm based on Markov distance, which is executed in the offline computing layer. Furthermore, the results are written into the user profile table and group profile table to support front-end visualization display and management decision-making.

2.3 U-Net improvement strategy and automatic acquisition of ground vegetation conditions

To further verify the applicability of the model in multi-scale feature extraction and complex vegetation distribution environments, the DeepLabv3+architecture module is introduced to enhance the U-Net based on the original U-Net model. This can improve its performance in automatic acquisition of ground vegetation conditions tasks (). The improved model maintains the advantages of the U-Net encoding decoding structure while integrating the dilated convolution of DeepLabv3+ with ASPP, achieving efficient perception of multi-scale features and full utilization of rich contextual information, as shown in Figure 5 ().

FIGURE 5

The optimized model introduces hollow convolution in the EP of U-Net, replacing the standard CO with a CK operation with dilation rate . The hollow convolution is shown in Equation 10.

In Equation 10, displays the value of the output FM at position , displays the input FM, displays the CK weight, displays the CK length, and displays the hole rate. When is present, the convolutional kernel expands the receptive field while keeping the number of parameters constant, enabling the model to simultaneously capture local texture features and large-scale structural features, especially suitable for artificial forest environments with significant changes in vegetation canopy scale (). The DeepLabv3+ specific ASPP () is added at the bottleneck of the EP. The ASPP module consists of multiple parallel hollow convolution branches, each of which uses different hole rates to perform multi-scale sampling on input features. Assuming the input feature is , ASPP outputs as shown in Equation 11.

In Equation 11, represents the convolution output with a hole rate of , and represents the globally average pooled feature. By concatenating these multi-scale features and performing -convolution fusion, the model can simultaneously utilize local details and global contextual information in a single forward propagation, effectively improving segmentation accuracy under complex terrain conditions. This mechanism is particularly critical for vegetation extraction, where targets exhibit significant spatial heterogeneity, ranging from individual tree crowns to extensive forest patches. Unlike standard convolutions with fixed receptive fields, the parallel branches of ASPP enable the network to capture fine-grained canopy textures and broad semantic distributions simultaneously. Consequently, this mitigates the issue of scale inconsistency, ensuring robust segmentation performance across varying vegetation growth stages and densities. In terms of clustering algorithm, improvements have been made on the basis of traditional K-means, introducing binary clustering strategy and Markov distance measurement. The traditional K-means algorithm clusters by minimizing the Euclidean distance between samples and cluster centers. However, in multidimensional feature spaces, when there are time series correlations or state transition features between features, the Euclidean distance cannot fully characterize the relationships between samples (). Specifically, Euclidean distance treats similarity relies solely on static geometric proximity, ignoring the directional and probabilistic nature of ecological changes. In contrast, Markov distance evaluates similarity based on the reachability between states within a diffusion process. Intuitively, this metric considers two samples similar not just because they are close in feature space, but because there is a high probability of transition between them, thereby effectively capturing the continuous dynamic evolution of vegetation health. Therefore, the study uses Markov chain models to characterize the probability of vegetation feature state transitions and measures sample similarity based on Markov distance. The definition of Markov distance is shown in Equation 12.

In Equation 12, represents the transition probability of sample in state , represents the transition probability of sample in the same state, and represents the total number of states. Through this measurement, the model can more sensitively identify the differences in vegetation status at different growth stages or under environmental disturbances (). The specific process of the binary K-means clustering algorithm is shown in Figure 6.

FIGURE 6

In Figure 6, the algorithm first considers all samples as a whole cluster. Determine whether further partitioning is necessary by calculating the variance and average variance of each sample feature. A sub-cluster is split into two sections when its variance exceeds the average variance, and the sample with the greatest and smallest distance is chosen to be the initial cluster centers, and . The cluster centers are iteratively updated until convergence. The algorithm uses Markov distance as a similarity criterion during each partition. This allows the clustering results to more accurately reflect the dynamic evolution of vegetation features. The final number of cluster centers, k, is adaptively determined by the internal variance distribution of the data rather than being preset, thereby improving the algorithm’s adaptability in complex ecosystems (). The proposed method is named the deep clustering model for automatic vegetation status recognition (DC-AVSR).

3 Performance analysis of vegetation information automatic acquisition method based on U-Net and clustering algorithm

3.1 Analysis of classification prediction effect

The RS data used in the experiment consists of GF-2 multispectral images with a spatial resolution of 3.24 m covering the Hebei Saihanba mechanized forest farm field. Cloud-free images are selected as the raw data during the vegetation growth season (June-September 2023–2024). The images undergo standard preprocessing steps, such as radiometric calibration, atmospheric correction, fusion, and cropping, prior to the experiment. Based on previous U-Net segmentation results, vegetation regions are extracted to provide feature input for subsequent clustering and visualization. The feature dimensions include normalized vegetation index and structural features, and are standardized uniformly. The experimental hardware platform is NVIDIA RTX3090 GPU (24 GB video memory), Intel XeonGold 6226R CPU (2.9 GHz, 32 cores), 256 GB memory, and the operating system is Ubuntu 20.04LTS. The DL framework adopts PyTorch 2.0, Python version 3.10, and is accelerated using CUDA 11.8. The study selects simple and efficient design for semantic segmentation with Transformers (SegFormer), local-enhanced multi-scale aggregation swin transformer for high-resolution RS semantic segmentation (LMA-Swin), attention-enhanced residual U-Net architecture (AER U-Net), dual-path attention residual U-Net for forest burned area detection (DPAttResU-Net), and the proposed DC-AVSR method for comparison. Every model is trained using the same partitioning for the training and validation sets. The input size is uniformly 256 × 256 pixels, the optimizer uses Adam, the initial learning rate (LR) is set to 1e4, and the batch size is 16. The comparison of classification prediction and clustering performance of different models on the reduced dimensional feature space is shown in Figure 7.

FIGURE 7

Figure 7a displays the clustering outcomes of DC-AVSR. The clear distribution boundaries, obvious inter cluster spacing, and high intra-cluster density indicate that the model has superiority in capturing differences in vegetation features. Figure 7b displays the outcomes of the SegFormer model. Although multiple clusters can be distinguished, there is overlap between some clusters, and the boundaries between clusters are not as clear as DC-AVSR. Figure 7c displays the outcomes of the LMA-Swin model. This model has shown some performance in multi-scale feature aggregation, but there is still a problem of inter cluster aliasing. Figure 7d displays the outcomes of the AERU-Net model. It can distinguish the main feature clusters, but the boundary clarity and clustering stability are slightly inferior to DC-AVSR. Figure 7e displays the clustering performance of the DPAttResU-Net model. Although it has a certain degree of aggregation, there is still a phenomenon of uneven distribution of transitional samples between different clusters. The study selects mean absolute error (MAE) and root mean square error (RMSE) for testing. Figure 8 displays the findings.

FIGURE 8

Figure 8a depicts the decrease in MAE of each model with the number of iterations. All models show a decreasing trend in error during the training process. However, the DC-AVSR model always maintains the lowest MAE at the same number of iterations and converges faster. This proves the superiority of the model in feature extraction and predicted accuracy. Figure 8b shows the trend of RMSE variation. The RMSE of different models also gradually decreases with the number of iterations, while the RMSE of the DC-AVSR model is always lower than that of other comparison models. Especially in the later stages of iteration, the error remains stable at a low level. This indicates that the model not only has advantages in overall error control, but also performs better in stability and robustness.

To verify the adaptability and generalization performance of the DC-AVSR model on multi-source and multi scene RS data, four typical public RS datasets are selected for the experiment: ISPRS 2D Semantic Labeling Contest–Vaihingen Dataset (ISPRS Vaihingen), DeepGlobe 2018 Land Cover Classification Challenge Dataset (DeepGlobe LCC), Gaofen Image Dataset for Land Cover Classification (GID), and BigEarthNet Large-Scale Sentinel-2 Benchmark Archive (BigEarthNet). Each dataset covers different scenes from high-resolution aerial images to medium resolution multispectral images, including urban vegetation, farmland, forests, and multi label ecological environment types. The data is uniformly preprocessed before the experiment, including radiometric calibration, atmospheric correction, image fusion, and standardization. All samples are cropped to 256 × 256 pixels and maintain a training, validation, and testing partition ratio of 8:1:1. The detailed characteristics of the datasets utilized in this study are summarized in Table 2.

TABLE 2

DatasetSensor/SourceSpatial resolutionSeason/Scene typeSample size (pixels)Split ratio (Train:Val:Test)
Saihanba (main)GF-2 PMS0.81 m (PAN)/3.24 m (MS)June–September (Forest)256 × 25608:01:01
ISPRS VaihingenAerialHigh-resolutionUrban Vegetation256 × 25608:01:01
DeepGlobe LCCSatelliteMedium-resolutionDiverse Land Cover256 × 25608:01:01
GIDGaofen-2High-resolutionLarge-scale Land Cover256 × 25608:01:01
BigEarthNetSentinel-2Medium-resolutionMulti-label Ecology256 × 25608:01:01

Summary of dataset characteristics and experimental settings.

The comparison of predicted accuracy and true accuracy of the DC-AVSR model on different publicly available RS datasets is shown in Figure 9.

FIGURE 9

Figure 9a reflects the performance of the model on the ISPRS Vaihingen dataset. As the sample set size increases, the accuracy difference between predicted values and true values gradually decreases. This indicates that the model has strong adaptability in high-resolution images of urban scenes. Figure 9b shows the results on the DeepGlobe LCC dataset. The overall trend of the model prediction curve is highly consistent with the true value, and tends to stabilize after the sample size reaches a certain scale. This demonstrates the robustness of the model to multiple types of land cover. Figure 9c presents the performance on the GID dataset. As the sample set increases, the predicted results quickly approach the true values and eventually almost overlap. This demonstrates the superior generalization ability (GA) of the model in large-scale high-resolution RSIs. Figure 9d shows the results of the BigEarthNet dataset. This model can maintain high consistency between predicted values and true values even under complex multi label features, and the predicted accuracy continues to improve as the sample set expands.

3.2 Analysis of fine-grained image segmentation effect

Further fine-grained analysis is conducted on the segmentation results. The experimental data is selected from GF-2 RSIs of the Hebei Saihanba mechanized forest farm area, with a resolution of 3.24 m, and obtained during the period of vigorous vegetation growth. All models are tested under the same training/validation/testing partition, with a uniform input size of 256 × 256 pixels. The Adam optimizer is used during the training process, with an initial LR of 1e 4. The cosine annealing scheduling strategy is used to dynamically adjust the LR. The evaluation of feature indicators includes MAE, RMSE, boundary refinement, feature overlap, average contour coefficient, etc.,. The performance of the model is comprehensively evaluated from three aspects: error control, boundary recognition, and feature consistency. The feature indicators of different models in fine-grained vegetation segmentation are shown in Table 3.

TABLE 3

ModelMAERMSESilhouetteBoundary accuracy (%)Overlap ratio (%)
DC-AVSR0.8121.0470.68493.2788.91
SegFormer0.8671.1230.65291.8485.36
LMA-Swin0.8941.1590.64190.2184.77
AER U-Net0.9231.1970.62289.6483.42
DPAttResU-Net0.9511.2440.60988.1582.03

Performance of different models in fine-grained vegetation segmentation based on feature indicators.

The overall trend in Table 3 shows that all models have MAE and RMSE indicators below 1.3, indicating that they have good accuracy in basic segmentation tasks. However, the DC-AVSR model achieves the lowest values in both indicators (MAE0.812, RMSE1.047), outperforming other models. At the same time, in terms of boundary refinement and feature overlap indicators, DC-AVSR reaches 93.27% and 88.91%, which is a significant improvement compared to the sub optimal model SegFormer’s 91.84% and 85.36%. The improved feature encoding and clustering strategy has stronger adaptability to boundary details and inter cluster differentiation. In addition, the average silhouette coefficient results also shows that DC-AVSR has better intra class consistency and inter class separability, further verifying its ability to capture subtle differences in vegetation status in high-dimensional feature spaces. The intra-cluster consistency and inter-cluster discrimination of the DC-AVSR model are tested at various k values in an attempt to examine the effect of clustering number on the segmentation performance of the model. Table 4 displays the findings.

TABLE 4

Number of clusters (k)Intra-cluster Dist.Inter-cluster Dist.Consistency (%)Separation (%)
30.2841.53791.7286.45
40.2671.58392.8487.36
50.2591.62193.4588.91
60.2511.60792.1388.27
70.2481.58891.7687.02

Segmentation consistency and differentiation of DC-AVSR model under different cluster numbers.

In Table 4, as the cluster increases from 3 to 5, the average intra-cluster distance of the model gradually decreases, while the inter cluster distance slightly increases. It displays stronger internal cohesion and external differentiation. When k = 5, the model reaches its optimal state. The consistency is 93.45% and the discrimination is 88.91%. At this point, feature space partitioning can accurately reflect the forest age gradient and health status, while avoiding excessive refinement caused by too many clusters. To establish an environmental interpretation for the optimal k = 5 clustering, these groups were explicitly mapped to documented biophysical criteria and field survey categories within the Saihanba forest farm. Specifically, the five clusters correspond to distinct ecological states: mature healthy forest with high canopy density, middle-aged actively growing forest, young planted forest with lower biomass, vegetation exhibiting early signs of environmental stress or diminished chlorophyll, and degraded or sparse vegetation areas. Aligning the algorithmic clustering with these tangible health status and forest age categories ensures that the extraction results provide interpretable and practical indicators for precision forestry management, thereby validating the biophysical relevance of the model beyond purely mathematical metrics. As the k value continues to increase to 6 or 7, there is a slight decrease in intra-cluster consistency and discriminability. Excessive clustering can introduce redundant categories and reduce the efficiency of the model in expressing ecological states.

3.3 Application effect of automatic vegetation information acquisition method

After the precise segmentation and clustering of vegetation features are extracted using the improved U-Net and DeepLabV3+ strategies, further research is conducted to apply the DC-AVSR model to multi-source, multi-scene RS data. This is done to evaluate the model’s adaptability and GA in the task of automatically acquiring vegetation information. Table 5 displays the outcomes of the experiment.

TABLE 5

Metric/DatasetISPRS VaihingenDeepGlobe LCCGIDBigEarthNet
Vegetation coverage (%)68.4274.1981.3677.58
Mean canopy width (m)3.272.943.122.88
Mean height (m)5.846.357.216.74
Chlorophyll index (CI)0.6830.7120.7450.728
Vegetation density index (VDI)0.7920.8150.8410.829
Classification accuracy (%)94.1394.8595.4694.72
FGS accuracy (%)88.2188.7989.3288.53
IoU0.870.880.890.88
F1 score0.920.920.930.92
Extraction speed (fps)24.825.125.425.0

Feature output performance of DC-AVSR model in multi dataset vegetation information automatic acquisition task.

In Table 5, DC-AVSR performs the most outstandingly on the GID dataset. The classification accuracy reached 95.46%, and the FGS accuracy reaches 89.32%. Meanwhile, the intersection over union (IoU) and F1 scores are 0.89 and 0.93, respectively, indicating that the model performs the best in segmentation and state recognition on high-resolution forest data. In addition, the model maintains high performance on multiple datasets, with an average extraction speed of about 25 frames per second. This indicates that it has good efficiency in real-time or near real time applications. The results of ecological feature extraction show that DC-AVSR has superior GA and stability in different scenarios. This provides effective technical support for large-scale monitoring and dynamic evaluation of vegetation health. The performance comparison of four models in automatically obtaining vegetation information at different sample ratios is shown in Figure 10.

FIGURE 10

Figure 10a illustrates the trend of accuracy of each model as a function of sample proportion. As the sample size increases, the accuracy of all models shows an upward trend. Among them, the DC-AVSR model maintains a leading position in almost the entire interval and significantly outperforms other models at high sample ratios. This indicates that it has stronger discriminative ability in the feature extraction and classification stages. Figure 10b shows the trend of recall rate changes. Although different models perform similarly at low sample ratios, as the data volume increases, the recall rate of DC-AVSR steadily improves and overall outperforms other models. This proves that it has better sensitivity in identifying vegetation boundaries and multi-scale features. Figure 10c shows the variation curve of F0.5 value. F0.5 focuses more on the contribution of accuracy. DC-AVSR consistently maintains a high level on this indicator. Especially in the stage of medium to high sample proportion, it shows obvious advantages, reflecting the model’s ability to achieve a balance between low false alarm rate and high accuracy in automated vegetation information acquisition. The comparison of IoU and running time between DC-AVSR and SegFormer models in different experimental areas is shown in Figure 11.

FIGURE 11

Figure 11a depicts the IoU performance of two models in different regions. The outcomes show that the DC-AVSR model is significantly better than the SegFormer model in all regions, with IoU values consistently approaching or exceeding 90%. This reflects its stronger modeling ability for boundary accuracy and regional consistency in fine-grained vegetation segmentation. Figure 11b shows a comparison of the running time of two models under the same task. Although the DC-AVSR model has a slightly longer running time when processing large area data, it still remains within an acceptable range. Meanwhile, it significantly outperforms SegFormer in terms of overall inference stability. The exclusion study of the proposed components on the GID dataset is shown in Table 6.

TABLE 6

Model configurationASPP moduleMarkov clusteringClassification accuracy (%)Intersection over union (IoU)
Standard U-Net (Baseline)89.150.78
U-Net + ASPP92.680.84
DC-AVSR (Proposed)95.460.89

Exclusion study of the proposed components on the GID dataset.

In Table 6, to isolate the individual contributions of the proposed components and evaluate their specific impacts, an exclusion study was conducted using the GID dataset. The original standard U-Net was utilized as the baseline model. Subsequently, the Atrous Spatial Pyramid Pooling (ASPP) module and the Markov distance-based binary K-means clustering algorithm were progressively integrated into the architecture. As shown in Table 6, the progressive improvement across the evaluation metrics explicitly verifies that each independent component systematically and significantly contributes to the overall performance enhancement of the proposed DC-AVSR method.

4 Conclusion

A DC-AVSR model was proposed and constructed to address the issues of relying on low-level features and low information quality in vegetation information acquisition for current agricultural and forestry RS monitoring. This method first adopted an improved U-Net network that integrated hollow convolution and ASPP in the DeepLabv3+ architecture to enhance the perception ability of multi-scale features. This achieved high-precision segmentation of vegetation and non vegetation areas in RSIs. Subsequently, to achieve FGS of vegetation internal states, an improved binary K-means clustering algorithm based on Markov distance was introduced. This algorithm could adaptively determine the number of clusters based on the internal variance of the data and effectively characterize the dynamic relationships between different vegetation states. The experimental results showed that compared with other advanced models, the MAE and RMSE of DC-AVSR reached the lowest levels of 0.812 and 1.047, respectively. In the FGS task, the boundary recognition accuracy of the model was 93.27%, and the average contour coefficient was 0.684. By optimizing the number of clusters, the model achieved optimal segmentation performance when the cluster center k = 5. The intra-cluster consistency and inter cluster discrimination were 93.45% and 88.91%, respectively. The application on the GID public dataset validated the effectiveness of the model, with a classification accuracy of 95.46% and a FGS accuracy of 89.32%. The IoU and F1 scores were 0.89 and 0.93, respectively, and the processing speed could reach 25.4 frames per second. Despite the promising results, several limitations remain to be addressed. First, the generalization capability of the model on alternative data sources, such as ultra-high-resolution UAV imagery or hyperspectral sensors, requires further validation, as the current training relies primarily on GF-2 satellite data. Second, the robustness of the algorithm against significant seasonal and phenological variations has not been fully explored, which is crucial for consistent year-round monitoring. Finally, the substantial computational demand of the deep clustering architecture poses challenges for deployment on resource-constrained edge devices. Future research will focus on integrating multi-source data to enhance dimensional perception, conducting long-term temporal analysis to adapt to phenological changes, and employing model compression techniques, such as pruning and quantization, to achieve lightweight and real-time processing capabilities. Furthermore, because the geographical focus is heavily concentrated on the managed temperate forest of Saihanba, the performance of the model in radically different ecosystems—such as tropical rainforests, savannas, or diverse agricultural croplands—remains unverified. Future research must incorporate heterogeneous datasets across various biomes and utilize sensors with diverse spectral characteristics, including Sentinel-2, to systematically evaluate and enhance the broad-scale applicability of the algorithm.

Statements

Data availability statement

The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.

Author contributions

HZ: Funding acquisition, Investigation, Methodology, Project administration, Resources, Writing – original draft, Writing – review and editing. MP: Investigation, Writing – original draft, Writing – review and editing.

Funding

The author(s) declared that financial support was not received for this work and/or its publication.

Conflict of interest

Author(s) HZ, and MP were employed by Alpha Earth Observation.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

  • 1

    AlahacoonN.EdirisingheM. (2022). A comprehensive assessment of remote sensing and traditional based drought monitoring indices at global and regional scale. Geomatics, Nat. Hazards Risk13 (1), 762799. 10.1080/19475705.2022.2045931

  • 2

    CaldersK.BredeB.NewnhamG.CulvenorD.ArmstonJ.BartholomeusH.et al (2023). StrucNet: a global network for automated vegetation structure monitoring. Remote Sens. Ecol. Conservation9 (5), 587598. 10.1002/rse2.322

  • 3

    ChenY.MortonD. C.RandersonJ. T. (2024). Remote sensing for wildfire monitoring: insights into burned area, emissions, and fire dynamics. One Earth7 (6), 10221028. 10.1016/j.oneear.2024.05.018

  • 4

    DashpurevB.DorjM.PhanT. N.BendixJ.LehnertL. W. (2023). Estimating fractional vegetation cover and aboveground biomass for land degradation assessment in eastern Mongolia steppe: combining ground vegetation data and remote sensing. Int. J. Remote Sens.44 (2), 452468. 10.1080/01431161.2022.2163456

  • 5

    ElairC.Rkha ChahamK.HadriA. (2023). Assessment of drought variability in the Marrakech-Safi region (Morocco) at different time scales using GIS and remote sensing. Water Supply23 (11), 45924624. 10.2166/ws.2023.238

  • 6

    GanjuN. K.CouvillionB. R.DefneZ.AckermanK. V. (2022). Development and application of landsat-based wetland vegetation cover and unvegetated-vegetated marsh ratio (UVVR) for the conterminous United States. Estuaries Coasts45 (7), 18611878. 10.1007/s12237-022-01081-x

  • 7

    HakamO.BaaliA.AzennoudK.LyazidiA.BourchachenM. (2023). Assessments of drought effects on plant production using satellite remote sensing technology, GIS and observed climate data in Northwest Morocco, case of the lower Sebou Basin. Int. J. Plant Prod.17 (2), 267282. 10.1007/s42106-023-00236-5

  • 8

    HamedianfarA.MohamedouC.KangasA.VauhkonenJ. (2022). Deep learning for forest inventory and planning: a critical review on the remote sensing approaches so far and prospects for further applications. Forestry95 (4), 451465. 10.1093/forestry/cpac002

  • 9

    HelaliJ.AsaadiS.JafarieT.HabibiM.SalimiS.MomenpourS. E.et al (2022). Drought monitoring and its effects on vegetation and water extent changes using remote sensing data in Urmia Lake watershed, Iran. J. Water Clim. Change13 (5), 21072128. 10.2166/wcc.2022.062

  • 10

    IppolitoM.De CaroD.CiraoloG.MinacapilliM.ProvenzanoG. (2023). Estimating crop coefficients and actual evapotranspiration in citrus orchards with sporadic cover weeds based on ground and remote sensing data. Irrigation Sci.41 (1), 522. 10.1007/s00271-022-00806-y

  • 11

    KarimiM.ShahediK.RazieiT.MiryaghoubzadehM. (2022). Meteorological and agricultural drought monitoring in Southwest of Iran using a remote sensing-based combined drought index. Stoch. Environ. Res. Risk Assess.36 (11), 37073724. 10.1007/s00477-022-02223-9

  • 12

    KumarV.SharmaK. V.PhamQ. B.SrivastavaA. K.BogireddyC.YadavS. M. (2024). Advancements in drought using remote sensing: assessing progress, overcoming challenges, and exploring future opportunities. Theor. Appl. Climatol.155 (6), 42514288. 10.1007/s00704-024-04917-y

  • 13

    KureelN.SarupJ.MatinS.GoswamiS.KureelK. (2022). Modelling vegetation health and stress using hyperspectral remote sensing data. Model. Earth Syst. Environ.8 (1), 733748. 10.1007/s40808-021-01124-x

  • 14

    LouX.HuangY.FangL.HuangS.GaoH.YangL.et al (2022). Measuring loblolly pine crowns with drone imagery through deep learning. J. For. Res.33 (1), 227238. 10.1007/s11676-021-01328-6

  • 15

    MarraE.LaschiA.FabianoF.FoderiC.NeriF.MastrolonardoG.et al (2022). Impacts of wood extraction on soil: assessing rutting and soil compaction caused by skidding and forwarding by means of traditional and innovative methods. Eur. J. For. Res.141 (1), 7186. 10.1007/s10342-021-01420-w

  • 16

    MadhaviM.KolikipoguR.PrabakarS.BanerjeeS.MaguluriL. P.RajG. B.et al (2024). Experimental evaluation of remote sensing-based climate change prediction using enhanced deep learning strategy. Remote Sens. Earth Syst. Sci.7 (4), 642656. 10.1007/s41976-023-00122-z

  • 17

    ManafifardM.HuangJ. (2025). A comprehensive review on wheat yield prediction based on remote sensing. Multimedia Tools Appl.84 (19), 2084320916. 10.1007/s11042-024-19253-x

  • 18

    MinyuW. A. N. G.YiL. U. O.QiaoyunX. I. E.XiaodanW. U.XuanlongM. A. (2022). Recent advances in remote sensing of vegetation phenology: retrieval algorithm and validation strategy. Natl. Remote Sens. Bull.26 (3), 431455. 10.11834/jrs.20211601

  • 19

    MoesingerL.ZottaR. M.van Der SchalieR.ScanlonT.de JeuR.DorigoW. (2022). Monitoring vegetation condition using microwave remote sensing: the standardized vegetation optical depth index (SVODI). Biogeosciences19 (21), 51075123. 10.5194/bg-19-5107-2022

  • 20

    MorganR. G.HodgsonM. E.WangC.SchillS. R. (2022). Unmanned aerial remote sensing of coastal vegetation: a review. Ann. GIS28 (3), 385399. 10.1080/19475683.2022.2078656

  • 21

    MullapudiA.VibhuteA. D.MaliS.PatilC. H. (2023). A review of agricultural drought assessment with remote sensing data: methods, issues, challenges and opportunities. Appl. Geomatics15 (1), 113. 10.1007/s12518-022-00481-9

  • 22

    OttoniA. L. C.NovoM. S. (2021). A deep learning approach to vegetation images recognition in buildings: a hyperparameter tuning case study. IEEE Lat. Am. Trans.19 (12), 20622070. 10.1109/TLA.2021.9480148

  • 23

    RoshaniS. H.RahamanM. H.RehmanS.MasroorM.AhmedR. (2023). Assessing forest health using remote sensing-based indicators and fuzzy analytic hierarchy process in Valmiki Tiger Reserve, India. Int. J. Environ. Sci. Technol.20 (8), 85798598. 10.1007/s13762-022-04531-1

  • 24

    RostamiA.Raeini-SarjazM.ChabokpourJ.ChadeeA. A. (2023). Soil moisture monitoring by downscaling of remote sensing products using LST/VI space derived from MODIS products. Water Supply23 (2), 688705. 10.2166/ws.2022.428

  • 25

    ThamagaK. H.DubeT.ShokoC. (2022). Advances in satellite remote sensing of the wetland ecosystems in Sub-Saharan Africa. Geocarto Int.37 (20), 58915913. 10.1080/10106049.2021.1963236

  • 26

    TuryahabweR.AndamaE.MulabbiA.NakiyembaA. (2024). Understanding the nexus between traditional brick-making, biophysical and socio-economic environment of Goma Division, Mukono Municipality, Central Uganda. J. Degraded Min. Lands Manag.11 (4), 63676378. 10.15243/jdmlm.2024.114.6367

  • 27

    WangY. P.ChangY. C.ShenY. (2022). Estimation of nitrogen status of paddy rice at vegetative phase using unmanned aerial vehicle based multispectral imagery. Precis. Agric.23 (1), 117. 10.1007/s11119-021-09823-w

  • 28

    YangD.MorrisonB. D.DavidsonK. J.LamourJ.LiQ.NelsonP. R.et al (2022). Remote sensing from unoccupied aerial systems: opportunities to enhance Arctic plant ecology in a changing climate. J. Ecol.110 (12), 28122835. 10.1111/1365-2745.13968

  • 29

    ZengY.HaoD.HueteA.DechantB.BerryJ.ChenJ. M.et al (2022). Optical vegetation indices for monitoring terrestrial ecosystems globally. Nat. Rev. Earth & Environ.3 (7), 477493. 10.1038/s43017-022-00298-5

  • 30

    ZhangS. H.HeL.DuanJ. Z.ZangS. L.YangT. C.SchulthessU. R. S.et al (2024). Aboveground wheat biomass estimation from a low-altitude UAV platform based on multimodal remote sensing data fusion with the introduction of terrain factors. Precis. Agric.25 (1), 119145. 10.1007/s11119-023-10082-x

Summary

Keywords

clustering algorithm, image segmentation, information acquisition, remote sensing, U-net, vegetation

Citation

Zhang H and Periyapperuma M (2026) Automatic acquisition method of ground vegetation conditions based on U-Net and clustering algorithm in remote sensing monitoring. Front. Environ. Sci. 14:1707781. doi: 10.3389/fenvs.2026.1707781

Received

29 September 2025

Revised

18 April 2026

Accepted

20 April 2026

Published

13 July 2026

Volume

14 - 2026

Edited by

Prashant K. Srivastava, Banaras Hindu University, India

Reviewed by

Abd Abrahim Mosslah, University of Anbar, Iraq

Harikesh Singh, University of the Sunshine Coast, Australia

Updates

Copyright

*Correspondence: Haoran Zhang,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics