ORIGINAL RESEARCH article

Front. Energy Res., 05 January 2026

Sec. Solar Energy

Volume 13 - 2025 | https://doi.org/10.3389/fenrg.2025.1691881

Short-term solar PV forecasting in microgrids using cloud top temperature and vision transformer based models

  • 1. Department of Energy and Climate Change, School of Environment, Resources and Development, Asian Institute of Technology, Pathum Thani, Thailand

  • 2. School of Electronic Engineering and Computer Science, Queen Mary University of London, London, United Kingdom

Abstract

Introduction:

Expanding clean-energy microgrids in remote areas is essential for achieving global decarbonisation and energy transition goals. Accurate short-term solar photovoltaic (PV) forecasting plays a key role in reducing diesel dependence, improving battery scheduling, and enabling reliable integration of renewable energy. However, forecasting remains challenging in many developing regions due to the lack of ground-based irradiance sensors, cloud cameras, and real-time monitoring infrastructure.

Methods:

This paper proposes a novel forecasting framework, termed CTT–ViT–Transformer, which integrates Generative AI techniques to enhance short-term solar PV forecasting in sensor-constrained microgrids. The framework employs Cloud Top Temperature (CTT) satellite imagery, capturing cloud height and thermal characteristics, processed through a Vision Transformer (ViT) for spatial feature extraction and a Transformer model for time-series prediction.

Results:

The proposed framework is evaluated using operational data from a real-world islanded microgrid. Results indicate that a standard Transformer model outperforms LSTM and CNN-LSTM baselines, achieving a mean absolute error (MAE) of 23.45 kW, root mean square error (RMSE) of 28.24 kW, and R² of 0.93. The CTT–ViT–Transformer further improves forecasting accuracy, reducing errors to an MAE of 15.99 kW and RMSE of 24.28 kW with an R² of 0.97, and consistently outperforms models relying on RGB satellite imagery. High predictive accuracy is maintained across four-step-ahead forecasts, with R² values exceeding 0.96.

Discussion:

The proposed approach requires no ground-based irradiance sensors, lowering adoption barriers for resource-constrained microgrids while remaining compatible with sensor-based data when available. Its scalability supports proactive energy management in the carbon-neutral microgrid on Koh Paluay Island by enabling more efficient scheduling of renewable generation and energy storage, thereby reducing fossil fuel use and operational costs. By enabling affordable and accurate forecasting, this framework aligns with Sustainable Development Goal 7 (Affordable and Clean Energy) and Sustainable Development Goal 13 (Climate Action), contributing to a just and sustainable global energy transition.

Graphical Abstract

Highlights

  • Introduces the first use of Cloud Top Temperature (CTT) for solar PV forecasting.

  • •Proposes a hybrid CTT–ViT–Transformer framework for forecasting in sensorless, remote microgrids.

  • •Achieves high accuracy (MAE 15.9868 kW, R2 0.9744), outperforming leading deep learning and Transformer models.

  • •Demonstrates the integration of Generative AI for solar PV output forecasting.

  • •Enables forecasting for microgrids or other generation sites without ground-based sensors or cameras.

1 Introduction

The urgency for clean energy transition continues to escalate as the global energy sector grapples with intensifying climate-related disruptions. Rising global temperatures, increasingly unpredictable weather, and more frequent natural disasters have placed unprecedented pressure on energy systems worldwide (Mischos et al., 2023). To address these environmental and systemic challenges, the United Nations adopted the 2030 Agenda for Sustainable Development in 2015. This comprehensive global framework advocates for progress that is inclusive, fair, and environmentally sustainable. At its core, Sustainable Development Goals (SDGs) 7 and 13 prioritise, respectively, the expansion of access to affordable clean energy and the reduction of greenhouse gas emissions to mitigate climate impacts.

In alignment with international commitments, Thailand has demonstrated notable progress in clean energy adoption and infrastructure modernisation. As of 2023, a national electrification rate of 99.99 percent had been achieved through the deployment of smart grid and microgrid technologies (Maidin, 2024). These technologies have facilitated broader electricity access while simultaneously reducing reliance on carbon-intensive fuels such as diesel and coal (Mohammadi et al., 2022). Despite these advancements, Thailand remains among the top global emitters, contributing approximately 0.86 percent of global emissions and ranking 19th worldwide. To address this, the Thai government has announced sector-specific climate goals, including a 57 percent reduction in energy-related carbon emissions by 2037 and a commitment to national carbon neutrality by 2050.

Achieving these targets requires system-level innovations, particularly in the management of decentralised renewable energy systems such as solar-powered microgrids. Within this context, accurate forecasting of solar photovoltaic (PV) power output plays a critical role in supporting grid reliability, optimising operational decisions, and facilitating renewable integration (Shams et al., 2021). The accuracy of forecasts is strongly dependent on both the reliability and accessibility of solar irradiance data, as these factors have a direct impact on photovoltaic system performance (). Improving forecast reliability is therefore essential not only for technical performance but also for long-term planning and investment in renewable infrastructure.

Among the various forecasting strategies related to solar PV power output, two principal approaches have emerged: direct and indirect forecasting methods. The direct approach estimates solar PV power output directly from historical generation data, avoiding the intermediate step of irradiance prediction. Deep learning (DL) techniques, particularly LSTM (Long Short-Term Memory) and CNN (Convolutional Neural Network) architectures, are commonly employed to capture temporal dependencies in the data while minimising reliance on physical system parameters (Konstantinou et al., 2021). This method has proven especially effective in settings where extensive historical power data are available ().

Comparative assessments have highlighted the superior performance of direct LSTM-based models in one-day-ahead forecasting scenarios. For example, normalised Root Mean Squared Error RMSE has been reported at approximately 25.6% for direct models compared with 28.1% for indirect approaches (Konstantinou et al., 2021). Further enhancements have been observed when synthetic irradiance data are incorporated, emphasising the potential of data augmentation techniques to boost forecasting performance ().

The indirect approach, by contrast, forecasts irradiance variables such as global horizontal irradiance (GHI) using meteorological or satellite datasets, subsequently converting these estimates into PV output via physical or data-driven models (). While this approach proves effective in contexts with abundant data, its predictive accuracy largely hinges on the resolution and accuracy of the irradiance inputs (Xu et al., 2025). Because the method involves two sequential steps, starting with irradiance forecasting and followed by power modelling, it tends to accumulate uncertainty throughout the process. This issue becomes particularly pronounced in areas with limited sensor infrastructure.

A key challenge in short-term solar power forecasting is the lack of reliable solar irradiance sensors at microgrid sites. Many locations either have no sensors or use low-grade instruments that require frequent calibration, leading to incomplete or inconsistent data (). Without accurate irradiance data, forecasting methods cannot be effectively applied, resulting in reduced accuracy for energy management and greater difficulty in long-term planning.

Research on irradiance forecasting has explored multiple approaches but still depends heavily on irradiance measurements. One study identified the optimal set of input variables for an artificial neural network (ANN) to predict monthly average global solar radiation, testing meteorological and geographical parameters to find the best-performing combination (). Another introduced the Bootstrap Aggregated Support Vector Machine, which outperformed standard Support Vector Machine (SVM) for hourly solar radiation forecasting with a very high correlation coefficient (R = 0.9913) ().

Furthermore, the hybrid model proposed by Gupta and Singh (2022) achieved a normalised mean absolute error (nMAE) of approximately 2.33% and a root mean square error (RMSE) of approximately 129.4 kW, outperforming baseline models and conventional methods. However, a key limitation is that it does not consider multi-step forecasting, which is essential for practical energy management in real-world systems. In addition, it relies heavily on ground-based meteorological measurements and employs a rule-based similar-day selection procedure, which reduces flexibility when dealing with highly variable weather conditions.

For the Solar Terms + Adaptive Boosting + Genetic Algorithm + Back Propagation (Adaboost–GA–BP) approach developed by Liu J. et al. (2023), the results showed an average RMSE reduction of 14.68% and an average MAE reduction of 23.42% compared with GA–BP and Adaboost–BP models, demonstrating superior performance in the context of northern China. However, major limitations include reliance on ground-based meteorological station data, making it unsuitable for applications in areas without comprehensive sensor infrastructure, and the fact that parameter tuning remains offline, which prevents the model from adapting in real time to data distribution changes.

Together, these studies underscore a single, critical limitation: their dependence on irradiance and sensor data. This reliance makes accurate solar PV output forecasting difficult in microgrids, particularly in remote or sensor-scarce areas, highlighting the urgent need for models that can perform reliably even when irradiance data are unavailable. To address this gap, recent research has turned to generative AI as a promising solution for mitigating data loss and sensor unavailability.

Currently, generative AI is being applied to address data loss and sensor failure problems. The adoption of models such as Wasserstein Generative Adversarial Networks (WGANs) has enabled the accurate reconstruction of missing PV data by generating realistic synthetic values that preserve temporal patterns (; Surathunmanun et al., 2024). Studies have shown that WGANs can reduce the mean squared error by more than 20%, even when up to 40% of the time series data are missing (Liu et al., 2025a; Zhang et al., 2020). In contrast, traditional imputation techniques often introduce bias and reduce model accuracy (Liu et al., 2025b).

Forecasting PV power output is increasingly complex because it depends on both solar irradiance and system characteristics, which has led many recent studies to adopt hybrid approaches. proposed the hybrid model, which combines data clustering, signal decomposition, and entropy analysis before forecasting with LSTM, reducing RMSE by 15%–20% but requiring extensive historical PV data. created an RF–CatBoost ensemble that achieved an R2 of 86% using a decade of weather and PV data, yet its reliance on large historical datasets limits its applicability in data-scarce microgrids.

A further study analysed a range of machine learning (ML) and deep learning (DL) models to forecast solar power generation, testing scenarios both with and without irradiance data. Interestingly, linear regression delivered the highest accuracy, with R2 = 0.99994. Although forecasting was possible without irradiance data, accuracy declined, reinforcing that irradiance is a critical variable. This study also highlighted that the tested models did not account for cloud movement or incorporate satellite imagery (Sharma et al., 2024).

Beyond individual models, broader review studies have examined the fundamental data sources and methods underpinning solar forecasting. Dheeraj et al. (2025) reviewed more than 200 studies from 2019 to 2023 and found that most solar forecasting methods rely on three main data sources: ground-based sky cameras, satellite imagery, and numerical weather indices. Sky cameras offer high spatial detail but are expensive to install and maintain and are difficult to deploy in remote microgrids.

Satellite imagery, particularly RGB products from platforms such as Himawari or Meteosat, provides wide coverage but only shows cloud shape and extent, not height or thickness, and requires substantial preprocessing. Many studies also use numeric data such as GHI or DNI, but these lack spatial structure and fail to capture real-time cloud dynamics. These limitations highlight ongoing constraints in both data quality and practical applicability for solar forecasting.

As a solution in cases where microgrids lack installed sensors, environmental variables such as temperature and humidity are often used as proxies. However, their forecasting performance is limited, especially in highly variable environments (Husein and Chung, 2019). An alternative approach involves integrating satellite imagery and numerical weather prediction models, although these methods often suffer from relatively low temporal and spatial resolution ().

Satellite imagery has therefore become an effective option in situations where ground-based solar irradiance data are unavailable or unreliable. Advanced platforms such as Himawari-8 and Himawari-9 provide continuous high-resolution atmospheric data, including cloud-top temperature, cloud motion, and multispectral imagery. These inputs are particularly valuable under heavy cloud cover, where surface irradiance readings are obscured (Xu et al., 2025).

According to the review by Husein et al. (2024), modern deep learning techniques are increasingly being applied to solar PV power forecasting, and Transformer-based models now form an important class of methods in this field. Within the family of Transformer-based architectures, three main lines of development can be identified. First, several studies focus on using Transformer models directly for PV power forecasting. Husein et al. (2024) employed a Transformer-based model for 30-min PV forecasting in the United Kingdom and achieved an R2 of 82.38%, while Husein et al. (2024) applied a Transformer model with multi-hour input horizons for hourly PV forecasting in Greece and reported an R2 of 92%. Second, hybrid models that combine Transformer components with convolutional networks have been proposed. For example, Husein et al. (2024) introduced a time-embedded Transformer–CNN hybrid model for hourly PV forecasting in Philadelphia and achieved an R2 of 87%.

In parallel with these architectural advances, a large body of research has emphasised multi-source data fusion and the use of synthetic weather-related data for PV forecasting. Ouyang et al. (2025) reported that integrating diverse data sources such as field sensor measurements, satellite imagery, and external indices can enhance forecasting performance but also introduces several practical challenges. Multi-source fusion often results in a large and redundant feature space, which increases computational burden and raises the risk of overfitting in the absence of appropriate feature selection or dimensionality reduction. Furthermore, the data processing pipeline must handle missing values and time synchronisation across data streams with different sampling frequencies and latencies. If these steps are not carefully designed, the physical relationships among key variables, such as solar irradiance, PV power output, and cloud structure, may be distorted (Ejiyi et al., 2025).

Data augmentation using synthetic weather-related data has also received considerable attention (Ouyang et al., 2025). Synthetic weather data can increase the number of training samples and extend coverage to a wider range of meteorological conditions, thereby mitigating overfitting when real observations are limited. However, such synthetic data are typically generated by recombining and reordering historical records rather than by fully simulating new atmospheric states, which limits their ability to represent genuinely novel patterns. In addition, the quality of synthetic data depends on the accuracy of the underlying models, so biases and systematic errors in the original forecasts may be propagated directly into the augmented datasets and subsequently into the PV forecasting models ().

Motivated by these considerations, this study adopts a Transformer-based forecasting framework to enhance prediction accuracy, combined with the use of satellite imagery as a surrogate for ground-based sensors. This design avoids the need for extensive multi-source data fusion, thereby reducing feature redundancy and computational burden, while also mitigating systematic errors that can arise from noisy or biased ground-based temperature observations.

The use of sky cameras has shown clear limitations in distinguishing between high, translucent clouds that partially transmit sunlight and low, dense clouds that almost completely block solar radiation. These cameras also tend to perform poorly under low-light conditions, such as during early morning hours or under heavily overcast skies (). Similarly, visible RGB satellite imagery is operationally constrained at night or before sunrise, leading to observation gaps when clouds form or move over solar power installations (Son et al., 2022). In addition, visible imagery cannot reliably provide information on cloud height or thickness, while the installation and maintenance of high-resolution sky imaging systems introduce additional costs and technical complexity ().

Cloud-top temperature (CTT) is a physically based variable derived from satellite thermal infrared signals. It provides a calibrated estimate of the temperature at the cloud top, which is directly related to cloud-top height and the vertical development of cloud systems, and indirectly linked to cloud optical thickness (Yirga et al., 2024). In recent studies, CTT data have emerged as an important source of information for characterising cloud height and thickness. Low CTT values indicate high, cold cloud tops that are typically tall and strongly convective, whereas higher CTT values tend to correspond to lower and potentially denser clouds. This property enables forecasting models to assess the shading impact of clouds on surface solar irradiance more effectively than when using visible imagery alone. Several studies have shown that infrared satellite indicators such as CTT exhibit strong correlations with ground-measured irradiance and outperform visible-spectrum satellite images (Son et al., 2022).

Nevertheless, the use of CTT in solar energy forecasting is still at an early stage. Originally, CTT was developed for meteorological applications, and only in recent years have high-quality geostationary satellites such as the GOES-R series and Himawari begun to provide high-resolution, high-frequency multispectral data, making the operational use of CTT for accurate forecasting increasingly feasible (). Low CTT values generally indicate tall, thick, and strongly convective cloud systems that significantly reduce the amount of solar radiation reaching the surface and thus cause substantial reductions in photovoltaic (PV) power output.

From a technical perspective, CTT is not an entirely new satellite product; rather, it is derived from standard thermal infrared channels. In this study, however, CTT is explicitly used as a physics-informed representation that is directly tied to cloud structures most relevant to PV power generation. This stands in contrast to visible cloud images, which only represent cloud reflectance and texture and whose appearance depends on the solar zenith angle and illumination conditions, causing image brightness to vary over time and making them unusable at night when there is no sunlight (Hammer et al., 2015). Similarly, raw thermal infrared imagery, although available both day and night, provides only brightness temperature, which can be influenced by the underlying surface and the intervening atmosphere. When clouds are thin or semi-transparent, the measured infrared temperature does not represent the true cloud-top temperature alone (Wang et al., 2024).

Solar PV output is highly sensitive to cloud characteristics, which directly influence the amount and quality of solar radiation reaching the Earth’s surface (Song et al., 2025). Clouds alter irradiance by absorbing, reflecting, and scattering sunlight, which affects both the direct and diffuse components. Cloud height is a critical factor: low-level clouds, typically found below 2 km with temperature differences (TD) ranging from 5 °C to 15 °C, significantly reduce PV output by blocking direct sunlight. In contrast, high-level clouds such as cirrus, located above 6 km with TD values greater than 40 °C, tend to have less impact and may even enhance diffuse irradiance under specific conditions ().

Cloud optical thickness and fractional coverage also influence irradiance levels. When cloud coverage exceeds 40%, surface irradiance begins to decline. At more than 80% coverage, especially when the Sun is low at approximately 30° above the horizon, irradiance can sharply decrease by nearly 900 W/m2. Conversely, thin cirrus clouds that do not fully block sunlight may increase irradiance by approximately 200 W/m2 due to enhanced scattering of diffuse light (Tzoumanikas et al., 2016). Taken together, these interactions underscore the importance of incorporating detailed cloud characteristics into solar forecasting models in order to improve accuracy under varying atmospheric conditions.

In this regard, CTT offers clear advantages over conventional satellite imagery for solar energy forecasting. To date, however, there have been no reported studies that employ CTT as the primary input variable for solar power forecasting in a microgrid context, making this line of work both novel and particularly relevant for islanded or remote microgrids without ground-based irradiance sensors, which must rely primarily on satellite data to characterise cloud conditions.

Another important challenge is that, although modern machine learning techniques can improve forecasting accuracy, they often require large datasets, are sensitive to missing data, and struggle to capture complex nonlinear relationships without detailed feature engineering (Tsai et al., 2023). These limitations can lead to substantial forecasting errors, particularly under highly variable weather conditions where prediction accuracy is critical to maintaining power system stability.

In addition, when dealing with highly dynamic, sequential image data, traditional models such as RNNs (Recurrent Neural Networks), GRUs (Gated Recurrent Units), LSTMs, and CNNs remain useful but are known to have limitations in processing long input sequences and adapting to real-time conditions (Tawn and Browell, 2022). Constraints in data transmission and computational capacity further hinder the timely retraining and updating of these models, especially in operational settings (; Islam et al., 2021).

As a result, the integration of satellite imagery with Vision Transformers (ViTs) represents a highly promising advancement in image-based solar forecasting. Originally developed for natural language processing and computer vision, Transformer architectures have recently been adapted for solar energy prediction with encouraging results (Mercier et al., 2024; ). Compared with conventional deep learning methods such as CNNs, ViTs can extract broad spatial features through self-attention mechanisms that capture large-scale cloud structures and complex spatial dependencies influencing surface irradiance ().

Early studies combining ViTs with recurrent components or meteorological data have reported superior forecasting accuracy over CNN and LSTM models, particularly under rapidly changing cloud conditions (). Studies have also shown that Transformers excel in temporal sequence modelling, enabling models to learn the dynamics of cloud evolution over time. For instance, Mercier et al. (2023) demonstrated that a Transformer-based framework improved 15-minute-ahead solar irradiance forecasting accuracy by more than 20% compared with traditional methods, owing to its ability to model long-range dependencies through self-attention mechanisms (Wang et al., 2024; Pospíchal et al., 2022).

Tian et al. (2022) further reported a reduction in mean squared error of approximately 0.04 kW and an improvement in mean absolute percentage error by 22%–29% when compared with GRU and DNN models. Additional evaluations showed R2 values as high as 0.9512, confirming the robustness of Transformer models under diverse weather conditions (Hu et al., 2024). Furthermore, recent findings indicate that ViTs outperform traditional CNNs and cloud segmentation techniques in both accuracy and generalisability when trained on sequential sky images (Liu Y. et al., 2023). The attention mechanism within ViTs enables strong spatial relationship modelling, capturing cloud shapes and movement patterns that are crucial for short-term PV forecasting.

On this basis, the present research proposes a ViT–Transformer framework that utilises CTT images processed by a ViT encoder to extract and analyse salient attention-based features. These features are combined with input data such as historical solar PV output and decomposed data corrected by a WGAN-based imputation model to address anomalies. The combined data then pass through a Transformer model that leverages attention mechanisms to enhance the accuracy and reliability of short-term solar PV power forecasting (Zhan et al., 2024).

In several related studies, CNN–LSTM architectures have been widely adopted as powerful hybrid models capable of capturing both spatial features through the CNN component and temporal dependencies through the LSTM layer (; ; Nahid et al., 2023; ). One of the key limitations of this class of models lies in the difficulty of capturing long-term dependencies due to the sequential nature of LSTM, which can lead to information loss over extended input sequences. By contrast, Transformer models can capture long-range dependencies more effectively without performance degradation as sequence length increases. They are also more efficient for large-scale datasets and more flexible in handling irregular or missing sequences ().

To improve prediction accuracy in microgrids that lack solar irradiance sensors, a generative AI (GenAI) agent is employed to optimise hyperparameters and significantly reduce training time. The overall framework is evaluated across four main dimensions:

  • -

    WGAN Performance in Solar PV Output Data Reconstruction: The WGAN successfully generated synthetic PV data by learning daily patterns. While the original dataset had an average output of 34.33 kW with a standard deviation of 22.95 kW,the synthetic data achieved a MAE of 15.25 kW and a slightly lower standard deviation of 20.74 kW. This closely matched the original data and effectively replaced corrupted records. The seamless integration of WGAN as a crucial component of the overall forecasting process is a key strength, significantly enhancing the framework’s robustness against incomplete data, a common issue in microgrids.

  • -

    Transformer Architecture for Solar Power Forecasting: The Transformer model consistently outperformed traditional models such as LSTM and CNN-LSTM. It achieved the lowest MAE (0.0023), RMSE (28.2390 kW), and highest R2 (0.9270). Specifically, it reduced MAE by 32.35% (compared to LSTM) and 30.30% (compared to CNN-LSTM), and increased R2 by 4.88% and 3.81%, respectively.

  • -

    Accuracy of CTT Satellite Imagery: CTT imagery yielded the highest accuracy in forecasting solar PV output. It achieved the lowest MAE (0.0016), RMSE (24.2782 kW), and highest R2 (0.9744). Compared to solar irradiance data, CTT reduced MAE by 30.43%, RMSE by 14.06%, and increased R2 by 5.11%. When compared to RGB satellite imagery, CTT reduced MAE by 11.11%, RMSE by 6.42%, and increased R2 by 1.22%.

  • -

    Efficiency of GenAI in Hyperparameter Tuning: Utilizing a GenAI agent to guide hyperparameter selection dramatically improved efficiency. It reduced training runs from 17,496 to 1,024 and cut computation time by 86.21% (from 376 h to 21.5 h), significantly decreasing training time without compromising result accuracy.

Leveraging GenAI, this study presents a highly effective and practical forecasting framework, the CTT-ViT-Transformer, designed for enhanced short-term solar PV output prediction in microgrids, especially those with limited or no sensors. Its key novelty points are:

  • -

    Pioneering Use of CTT Satellite Imagery as a Primary Input: There is currently no published research that directly utilizes CTT satellite imagery as a primary input for solar PV output forecasting. This study demonstrates that CTT, which provides deep insights into cloud height and density, yields significantly better results than solar irradiance data or standard RGB satellite imagery. This directly translates to a marked improvement in solar PV output forecasting accuracy.

  • -

    Novel Hybrid Model CTT-ViT-Transformer: This work introduces the first integration of a ViT for processing CTT satellite imagery to extract its distinctive spatial features. These CTT-derived features, combined with other pertinent variables, are subsequently fed into a Transformer Model optimized for time series analysis to model the temporal relationships of solar PV output. Crucially, this architecture facilitates highly accurate and reliable forecasting that operates effectively without the need for solar irradiance or any other sensors, consistently outperforming models relying on conventional satellite imagery or even those with irradiance sensor inputs.

In summary, this framework addresses the critical challenge of solar PV forecasting in sensor-limited settings. It was applied and validated on a real-world remote solar microgrid on Koh Paluay Island, Surat Thani Province, Thailand, which lacks irradiance and other ground sensors and targets carbon neutrality. This application demonstrates substantial gains in prediction accuracy, enabling robust energy management aligned with carbon neutral objectives.

The remainder of this paper is structured as follows. Section 2 presents the research methodology and the proposed framework, including an overview of the Koh Paluay microgrid, the solar PV output data, temporal feature engineering, the use of cloud-top temperature satellite imagery, feature extraction with correlation analysis, and the Transformer-based solar PV forecasting model. Section 3 introduces the performance evaluation metrics. Section 4 presents and discusses the results. Section 5 concludes the study and outlines directions for future research.

2 Methodology and proposed framework

This paper outlines the complete framework for forecasting solar PV output employing the CTT-ViT-Transformer approach, as shown in Figure 1. The process starts with data acquisition and progresses to practical model implementation. Generative AI is incorporated at four critical phases:

FIGURE 1

The first component focuses on improving data quality by using WGANs to generate synthetic data to replace damaged or missing parts. This helps fill in missing sensor inputs and enhances model performance, particularly in data-limited environments.

The second component applies ViTs to extract spatial features from satellite images, capturing critical factors. ViTs leverage self-attention and dimensionality reduction to identify important patterns without relying on ground-based sensors.

The third component introduces Transformer architectures for time-series forecasting. Transformers use parallel attention mechanisms to process multiple variables, improving forecasting accuracy and adaptability to changing conditions.

The final component addresses hyperparameter optimization. A GenAI agent defines focused parameter ranges, enabling efficient grid search to tune variables such as learning rate and model depth. Early stopping is also employed to reduce training time once performance stabilizes.

2.1 Microgrid in Koh Paluay, Surat Thani, Thailand

This study examines the microgrid system on Koh Paluay, an off-grid island located in Surat Thani Province, Thailand. The island previously relied primarily on diesel generators, which led to environmental pollution and annual carbon emissions of approximately 537 tons. Koh Paluay comprises 228 households and 6 businesses, with a total installed capacity of 226.05 kW and an effective power output of around 156.74 kW. By 2024, electricity demand is projected to rise to 0.29 MW, while the island’s solar energy potential is estimated at 1.46 million kWh per year (Provincial Electricity Authority, 2024).

According to the data, if residents switch from using their personal diesel generators to electricity provided by the Provincial Electricity Authority (PEA), it will result in diesel generation occurring only during periods when solar energy is uncertain—particularly during the monsoon season. A critical issue is the shortage of solar irradiance sensors, which significantly affects microgrid management, particularly in forecasting solar PV output. This presents a major challenge to achieving the primary goal of establishing a carbon-neutral microgrid.

This site was chosen as a national model for a carbon-neutral microgrid in Thailand. Its infrastructure features 1,000 kWp of solar PV panels, two 300 kW diesel generators, and a battery energy storage system with a rated power of 750 kW and a total energy capacity of 1,500 kWh. You can see this visually represented in Figure 2, with a detailed facility single-line diagram in Figure 3.

FIGURE 2

FIGURE 3

Koh Paluay faces several real-world energy challenges, including limited energy access and rising demand. Crucially, the microgrid operates without irradiance sensors, which significantly hinders the implementation of AI-based forecasting and advanced energy management strategies. Currently, the system relies on manual, rule-based control for switching between diesel and battery sources. This lack of adaptability often leads to inefficiencies, particularly when dealing with sudden or unpredictable changes in energy data.

Therefore, this research has selected the Koh Paluay microgrid as a scalable and replicable model for sustainable energy management in remote regions, offering valuable insights for similar contexts.

2.2 Data acquisition and preparation

This section outlines the key processes involved in data collection and preparation, as detailed below.

2.2.1 Solar PV output – Koh Paluay microgrid

An analysis of solar PV output data was conducted for the Koh Paluay microgrid using 15-min interval data collected over a 1-year period. Historical data was recorded from 7 May 2024 to 30 April 2025, at 15-min intervals between 05:00 and 19:00, across 53 measurement points, as illustrated in Figure 4. This timeframe corresponds to the period since the official operation of the Koh Paluay microgrid began.

FIGURE 4

During the data collection period, some intervals were affected by maintenance activities in the microgrid, resulting in partial data loss. In this study, missing or corrupted data was corrected using the WGAN, applied specifically during periods involving maintenance or technical issues. The WGAN model was trained on daily solar PV generation patterns and was used to generate synthetic data to replace the missing segments during those periods.

WGAN is based on the Generative Adversarial Network (GAN), originally introduced by . GANs are a machine learning framework inspired by minimax game theory that enables adversarial learning. WGAN consists of two main components: a generator, which learns the distribution of real data to produce synthetic samples, and a discriminator, which evaluates whether a given sample is real or generated.

To address these limitations, WGAN was developed. This variant replaces the original divergence function with the Wasserstein Distance, resulting in a more stable training process. As a result, WGAN improves model reliability and mitigates several common weaknesses associated with conventional GANs.

The WGAN loss function uses the Wasserstein Distance because it provides more stable gradient updates.

The WGAN equation is:

Where

  • -

    is the Wasserstein Distance between the real data distribution and the generated data distribution

  • -

    is a joint distribution over real and generated data pairs.

Implementation, WGAN replaces the traditional discriminator with a network called the Critic, which is responsible for approximating the Wasserstein Distance between real and generated data. The corresponding loss function used in WGAN is defined as follows:

Where

  • -

    is the Critic’s output for real data .

  • -

    is the Critic’s output for generated data

WGAN outperforms conventional GAN in reducing mode collapse and producing more diverse outputs. Its robust architecture has enabled successful applications across various domains, including synthetic data generation, high-quality image and video synthesis, and renewable energy forecasting in microgrids (Mansour et al., 2024; Sun et al., 2023).

To address the issue of missing solar PV output data, WGAN was employed as a data reconstruction solution. Approximately 5.3% of the data was found to be missing due to sensor malfunctions, communication failures, or hardware issues. The original dataset recorded a minimum value of 0 kW, a maximum of 617.24 kW, an average of 34.33 kW, and a standard deviation of 22.95 kW.

WGAN was used to generate synthetic data and reconstruct complete daily sequences. The WGAN-generated data achieved a MAE of 15.25 kW relative to the original values and exhibited a slightly lower standard deviation of 20.74 kW, as shown in Figure 5. These results indicate that the generated data follow a stable distribution closely aligned with the original dataset, supporting their suitability for downstream forecasting applications.

FIGURE 5

2.2.2 Temporal feature engineering

After imputing missing solar PV values, the solar PV time series was decomposed into trend, seasonal, and residual components and then normalized (Nahid et al., 2023). Cyclical time encodings were added via sine and cosine of the clock time (SinTime, CosTime) to expose diurnal periodicity to the learner. The preprocessing pipeline is summarized in Figure 6 (Time-Series Decomposition of Solar PV Output into Trend, Seasonal, and Residual Components).

FIGURE 6

Pearson correlation analysis was performed to explore the associations between the decomposed components and the temporal encoding variables, as illustrated in Figure 7. The findings reveal that total solar PV output has a strong positive correlation with the seasonal component (r = 0.66).

FIGURE 7

Additionally, a moderate negative correlation was observed with the CosTime variable (r = −0.34), while a weak positive correlation was found with SinTime (r = 0.22). These findings highlight the influence of temporal patterns captured through time encoding, supporting the inclusion of these variables in the subsequent forecasting model.

In this paper, the variables used for comparison are those commonly applied in solar irradiance forecasting research. These include irradiance parameters such as DHI, GHI, and GTI, which are frequently sourced from Solcast due to their proven accuracy. The emergence of high-quality datasets derived from satellite observations has opened up new possibilities for improving the precision of solar PV output forecasts.

Solcast’s satellite-derived data, such as GHI, has been validated against ground-based measurements, with results confirming its high quality and international reliability. This makes it well-suited for academic research and solar energy potential assessment (; Mabasa et al., 2022).

A clear outcome is observed when comparing satellite-based data with ground truth: the relative Root Mean Square Error (rRMSE) of Solcast’s GHI is approximately 3.4% under clear-sky conditions and around 25.6% during overcast periods (). These results highlight that while Solcast provides high accuracy under stable weather conditions, its performance tends to decline in regions characterized by high variability or persistent cloud cover. CTT Satellite Imagery: ViT-Based Spatial Feature Extraction.

2.2.3 Cloud top temperature satellite imagery

To capture these cloud characteristics in a physically meaningful way, this study uses satellite-derived CTT imagery as a primary input. Primary CTT data are obtained from the Himawari-9 geostationary satellite. A Python script was developed to extract satellite images over Koh Paluay, Thailand, at 15-min intervals. Geostationary satellite frames are cropped over the island, with modalities including CTT and, where available, RGB composites. Each frame is then resized to 224 × 224 pixels to standardize ground sampling before being processed by a Vision Transformer (ViT) model to identify salient regions within each frame.

ViTs, proposed by , repurpose the Transformer framework for image processing by representing an image as an ordered sequence of patches that are processed through self-attention mechanisms. ViTs have demonstrated promising results in analyzing weather and cloud-related imagery. For instance, applied ViTs to classify weather conditions from ground-based camera images, distinguishing among 11 different categories (e.g., sunny, cloudy, foggy) with an accuracy of 93.5%. In the context of satellite-based cloud segmentation, Zhang et al. (2023) proposed CloudViT, a lightweight ViT architecture designed to detect cloudy pixels in satellite images.

ViTs are also effective in forecasting applications. Mercier et al. (2024) introduced a ViT model that predicts global and diffuse solar irradiance using all-sky camera images. The model is trained to focus on pertinent visual features, including elements such as cloud cover and the position of the sun. In coastal meteorology, Kamangir et al. (2024) developed FogNet-v2, a physics-informed ViT that predicts 24-h fog presence by tokenizing both image and physical features and using attention to capture fog-related patterns.

Liu J. et al. (2023) introduced a multimodal framework built on Transformer architecture to forecast ultra-short-term solar irradiance by leveraging time-sequenced sky imagery. The rapid movement of clouds often causes sudden changes in irradiance, making traditional methods less effective. Their approach uses ViTs to extract spatial features and capture global and long-range dependencies, which leads to improved forecasting accuracy over CNN-based and traditional cloud detection methods.

To interpret ViT predictions, researchers often analyze attention weights, particularly those from the special class token to the image patches. While raw attention maps can highlight areas of interest, they are not always reliable indicators of model reasoning (Jain and Wallace, 2019). To address this limitation, various methods have been devised to consolidate and enhance attention signals, yielding more informative saliency maps. These include attention rollout () and relevance propagation (), both of which enhance interpretability in ViT-based models.

The ViT architecture starts by segmenting an image into uniformly sized patches, which are subsequently flattened into vectors and projected via a linear embedding layer. These embedded patches, augmented with positional encodings, are then input into a Transformer encoder that jointly processes the entire set of patches. This design enables the model to capture global context effectively, making it suitable for complex visual tasks ().

Building on these capabilities, the present study focuses on extracting attention maps from satellite-derived cloud-top temperature images and incorporating them into a Transformer-based learning framework, along with other relevant variables. The image attention extraction process is illustrated in Figure 8 and consists of several sequential steps.

FIGURE 8

This paper adopts the base ViT architecture with the backbone “google/vit-base-patch16-224,” initialized from pretrained weights and run in evaluation mode so that all parameters remain as in the reference checkpoint. Input images are resized to 224 × 224 pixels in RGB; for grayscale or CTT inputs, the single channel is replicated to three channels before ingestion. Each image is partitioned into 16 × 16-pixel patches, yielding a 14 × 14 grid (196 patches). A sequence of 197 tokens is formed, comprising one classification token and 196 patch tokens. Each token is projected to a D = 768 embedding and augmented with learnable positional embeddings to preserve spatial context, then processed by a Transformer encoder with 12 blocks. Each block employs 12-head self-attention followed by a feed-forward network with an intermediate dimension of 3072 to enhance nonlinearity and feature mixing. Dropout and attention dropout are kept at the checkpoint defaults without modification to preserve the pretrained behaviour and reduce unnecessary tuning during evaluation on the datasets used in this work.

The self-attention identifies the five patches receiving the highest attention. Their spatial coordinates (x, y) are extracted, and Euclidean distances relative to the image center are computed. These derived features are then integrated as auxiliary inputs to the downstream forecasting model.

2.2.4 Feature extraction and correlation analysis

After applying the Vision Transformer (ViT), the five most attended regions are identified for each frame. As shown in Figure 9, the left panel displays the original satellite crop, whereas the right panel shows the ViT attention heatmap. For RGB imagery, attention concentrates on dense cloud structures distributed across the scene. For cloud-top-temperature (CTT) imagery, the attention map more closely follows cloud motion and separates clouds at different altitudes. The five most attended patches are highlighted by red rectangles. For each patch, the pixel-space coordinates (x,y) and the Euclidean distance to the image center are extracted.

FIGURE 9

Beyond decomposition into trend, seasonal, and residual components to isolate diurnal regularity and transient disturbances, cyclical time encodings are added using the sine and cosine of clock time (SinTime and CosTime). Additional scene-level features are then derived directly from the imagery: the mean grayscale intensity; the mean edge magnitude from Sobel gradients; the fractional cloud-cover ratio via scene-adaptive thresholding; and CTT-based statistics, including the image-wide mean and the standard deviation of the CTT-derived grayscale. Within the attended regions, the mean edge strength and the mean CTT are further summarized to assess whether attention concentrates on sharply delineated cloud boundaries and on higher, colder convective clouds or, instead, on lower stratiform layers. The complete set of variables is reported in Table 1 and is designed to link physically meaningful cloud properties and model attention patterns to short-term variability in solar PV output.

TABLE 1

VariableMeaning
Gray_scaleMean grayscale of the image
Edge_strengthMean edge/gradient magnitude
Xtop_k, Ytop_kPixel coordinates of the Top-5 ViT attention patches; indicate where the model focuses most
Distance_kNormalized euclidean distance from image center to each attention patch
CTT_meanMean cloud-top temperature from satellite; inversely related to cloud-top height/thickness

New Variable definitions used in the correlation analysis.

For each of the top-5 salient patches , the following descriptors are computed and used as auxiliary variables:

Where , denotes the number of patches per row/column.

Attention-weighted luminance (Gray%). For RGB inputs, luminance per pixel () is computed via Rec. 601 coefficients:

The attention-weighted mean luminance is:where is the normalized attention weight of patch . For CTT inputs, normalized brightness values are used with

The mean edge-magnitude of the image is obtained using the Sobel operator:where and are horizontal and vertical image gradients computed with the 3 × 3 Sobel operator. The square-root term is the gradient magnitude at each pixel. The spatial mean provides a global measure of edge strength in the scene, reflecting the sharpness of cloud boundaries and the prevalence of fine-scale structure.

The Euclidean distance of patch from the image center is normalized as:

The spatial variability of CTT-derived brightness is expressed by the standard deviation of its grayscale intensity:where is the mean grayscale intensity of the entire image, computed as the arithmetic average of over all pixels; and are the image width and height (in pixels), and is the grayscale intensity at pixel .

The cloud coverage is defined as:where 1{⋅} is the indicator function that equals 1 if the condition is true and 0 otherwise, and is a grayscale threshold defining cloud pixels. The metric returns the fraction of pixels classified as cloud. The choice of should be documented, for example fixed, scene-adaptive, or calibrated with solar geometry and sensor characteristics.

The scene-average cloud-top temperature is computed as:

Subsequently, Pearson correlations were computed between the ViT-derived descriptors and the solar/irradiance variables. Using RGB inputs (Figure 10), both the top-5 saliency descriptors and the distance-based summaries exhibit strong associations with PV output as well as DHI, GHI, and GTI. Using CTT inputs (Figure 11), the associations are generally stronger, and the salient regions span a broader and more physically coherent spatial pattern, thereby improving interpretability of solar attenuation by high-level clouds versus low-level clouds.

FIGURE 10

FIGURE 11

This paper also reports a correlation analysis (Table 2) using attention-based features from ViT top-attention regions (Figure 12) to relate cloud properties, irradiance, and PV output. The scene-average CTT_mean is moderately positively associated with PV (r = 0.456), consistent with warmer CTT indicating lower or thinner clouds and greater irradiance. In contrast, structural variability is negatively related to PV: CTT_gray_SD (r = −0.657), Edge_strength (r = −0.686), and Cloud_cover_ratio (r = −0.255). Within salient regions, Attn_edge_mean correlates negatively with PV (r = −0.402) and with DHI, DNI, and GTI (≈−0.37); Attn_cloud_height_mean shows negative correlations with PV (r = −0.359) and with DHI, DNI, and GTI (≈−0.26 to −0.27). Attn_cloud_height_mean is strongly negatively correlated with CTT_mean (r = −0.873) and positively with Cloud_cover_ratio (r = 0.829). Overall, attention-based features compactly capture cloud signals that attenuate surface irradiance and reduce Solar PV output.

TABLE 2

VariableMeaning
CTT_meanScene-mean cloud-top temperature; higher values typically imply lower/thinner clouds or clearer sky
CTT_gray_SDStandard deviation of grayscale over the whole image; proxy for spatial variability/texture of the cloud field
Cloud_cover_ratioFraction of image area above a brightness threshold; crude proxy for cloud coverage
Edge_strengthProminence of edges/texture across the full frame (indicates broken cloud boundaries)
Attn_edge_meanMean edge prominence within the 5 top attention patches (model-salient regions)
Attn_cloud_height_meanMean cloud-top height over the 5 top attention patches, derived from per-patch CTT

New Variable definitions used in the correlation analysis.

FIGURE 12

2.2.5 Forecasting solar PV output with transformer

Once the cloud top temperature imagery and associated features were gathered as outlined, the dataset was partitioned into 70% for training, 15% for validation, and 15% for testing purposes. The Transformer model was subsequently employed to perform the forecasting.

The Transformer, introduced by Vaswani et al. (2017), processes sequential data using self-attention mechanisms rather than relying on recurrent structures such as RNNs or LSTMs. Its architecture allows the entire input sequence to be processed in parallel, significantly improving both efficiency and accuracy across a variety of tasks, including machine translation, text generation, and time-series forecasting (Kalyan, 2024).

Figure 13 illustrates the Transformer’s encoder-decoder architecture. The encoder converts the input sequence into hidden representations through multi-head self-attention and feedforward networks. Subsequently, the decoder produces the output by attending to the encoder’s hidden states alongside earlier generated tokens. This mechanism enables the model to capture intricate temporal dependencies, thereby enhancing forecasting accuracy.

FIGURE 13

A core innovation is the attention mechanism, which assigns each token a query, key, and value to compute attention scores. Multi-head attention enables the model to capture diverse contextual dependencies in parallel, enhancing its ability to learn long-range relationships ().

This architecture makes the Transformer highly effective for sequence modeling across various domains beyond NLP, including forecasting and image analysis. Recent work has shown that large language models (LLMs), when guided by carefully designed prompts, can provide effective initial settings for machine learning hyperparameters and reduce manual trial-and-error (Tao et al., 2024). LLM-based assistants have also been reported to lower the overhead of search design while clarifying trade-offs and performance evaluation practices (Yao et al., 2025). In parallel, schema-constrained prompting, that is, forcing the model to respond in a specified JSON format with bounded parameter ranges, has been shown to reduce out-of-range outputs, off-topic responses, and hallucinations (Li et al., 2024). In addition, the Belief–Desire–Intention (BDI) paradigm, originally proposed by Bratman and later formalised for rational software agents, provides a structured way to elicit goal-oriented, explainable decisions from such systems (Ciatto et al., 2025).

Building on these ideas, this study applies prompt engineering and a BDI-inspired specification to employ a GenAI assistant, namely OpenAI’s GPT-4.1 model (temperature = 1.0, top-p = 1.0), exclusively in the initial stage of hyperparameter tuning. The role of GenAI is limited to proposing initial search ranges for grid search, consistent with the existing Transformer time-series literature, rather than directly tuning the model. The main objectives are: (i) to reduce the time spent guessing broad hyperparameter ranges over multiple iterations, (ii) to avoid excessively wide, unfocused search spaces, and (iii) to lower the computational burden of grid search by focusing on ranges that are plausible both in prior studies and under the hardware constraints of the target microgrid system.

The GenAI assistant operates under hard constraints derived from the literature (for example, bounds for learning rate, dropout, number of layers, and attention heads) and is instructed to refine ranges rather than freely invent them. It receives only high-level metadata such as dataset statistics, hardware limits, and a short summary of a few preliminary runs, and has no access to the validation or test sets, which are reserved strictly for subsequent grid search in order to avoid overfitting to evaluation data. If a GenAI suggestion falls outside the predefined bounds or violates the required JSON schema, the system automatically reverts to conservative, literature-based default ranges and flags the suggestion with a warning for manual review. Prompts are structured according to the BDI framework so that the assistant must justify each proposed range with explicit reasons that link back to prior work and microgrid-specific constraints.

Once a satisfactory set of search ranges has been obtained, the remaining procedure follows a conventional grid-search pipeline without further GenAI involvement. Models are trained on the training set and evaluated on a held-out validation set. The configuration achieving the lowest validation RMSE is then retrained once and evaluated on the test set.

3 Model’s performance evaluation metrics

The predictive accuracy of the models developed in this study is assessed using the test dataset. The forecasted outputs are visualized and analyzed with established statistical measures to enable an in-depth comparison between the CTT-ViT-Transformer and the baseline approaches. To ensure comprehensive and consistent evaluation, the following statistical metrics are employed in this study:

The MAE quantifies the average magnitude of prediction errors without considering their direction.

  • -

    Root Mean Square Error (RMSE) ()

RMSE measures the square root of the average squared differences between predicted and observed values, emphasizing larger errors.

  • -

    Pearson’s Correlation Coefficient (r) ()

This coefficient assesses the linear correlation between the actual and forecasted outputs.

  • -

    Coefficient of Determination (R2) ()

R2 indicates the proportion of variance in the observed data explained by the model.

4 Discussion

This paper evaluates the impact of three imputation methods (WGAN-based, mean, and linear interpolation) on Transformer-based solar PV forecasting. As shown in Table 3, the WGAN-imputed dataset yields the best performance, with the lowest MAE (23.4518) and RMSE (28.2390 kW) and the highest R2 (0.9270). Mean imputation gives the poorest results, with a MAE of 27.5485, an RMSE of 30.0015 kW, and an R2 of 0.8768, while linear interpolation performs slightly better than mean but still worse than WGAN, with a MAE of 24.9352, an RMSE of 28.8743 kW, and an R2 of 0.9227. These results indicate that the WGAN-based method provides more informative and physically consistent imputations for training the Transformer model.

TABLE 3

Imputation methodMAE (kW)RMSE (kW)R2
WGAN23.451828.23900.9270
Mean27.548530.00150.8768
Linear24.935228.87430.9227

Performance of Transformer models trained on datasets imputed with WGAN, mean, and linear interpolation.

As shown in Figure 14, the time series of PV power output cleaned and imputed using the WGAN, mean, and linear methods are compared. The WGAN-based imputation preserves the diurnal characteristics more faithfully, whereas the mean and linear methods produce unrealistic values in parts of the series, which is consistent with the performance of the Transformer models trained on these datasets reported in Table 3.

FIGURE 14

The forecasting performance of three deep learning architectures was assessed for predicting solar PV output in the Koh Paluay microgrid. The models compared were the Transformer, LSTM, and CNN-LSTM. As summarized in Table 4, the Transformer outperformed the other approaches, delivering the lowest MAE of 23.4518 and an RMSE of 28.2390 kW, together with the highest R2 value of 0.9270, reflecting a strong correspondence between observed and predicted outputs. By comparison, the LSTM and CNN-LSTM models achieved slightly lower accuracy in their predictions.

TABLE 4

ModelMAE (kW)RMSE (kW)R2
Transformer23.451828.23900.9270
LSTM28.342532.09370.8839
CNN-LSTM29.003333.05010.8929

Performance comparison of deep learning for solar PV output forecasting.

When benchmarked against external studies, such as the LSTM-based approach by , the proposed Transformer model delivered comparable results, despite differences in input features and operational contexts. These findings confirm that the Transformer model provides the most accurate prediction performance among the models evaluated in this study.

The evaluation of different input types demonstrates that CTT satellite imagery is the most effective standalone input for forecasting solar PV output. When used individually, CTT imagery achieved a MAE of 15.9868, an RMSE of 24.2782 kW, and an R2 of 0.9744, which outperforms solar irradiance data and RGB imagery as shown in Table 5.

TABLE 5

Input data typeMAE (kW)RMSE (kW)R2
Solar irradiance data23.451828.23900.9270
Satellite imagery (RGB)17.194525.94240.9626
CTT satellite imagery15.986824.27820.9744
Combined: Irradiance data, satellite imagery (RGB), and CTT satellite imagery15.854424.24450.9786

Comparison of input data types for forecasting solar PV output.

The advantage of CTT data lies in its ability to capture vertical cloud characteristics such as height, thickness, and thermal structure. These features are essential for accurately modeling surface-level solar irradiance. Unlike RGB imagery, which provides only visual cloud coverage, CTT imagery offers thermal information that allows the model to better anticipate fluctuations in solar radiation. This makes it a suitable alternative to ground-based irradiance sensors, especially in remote areas where real-time sensor data may be limited or unavailable.

When combining all three data sources, including irradiance data, RGB imagery, and CTT imagery, the model produces the best performance. This configuration results in a MAE of 0.0015, an RMSE of 24.2445 kW, and an R2 of 0.9786. The results confirm that integrating multiple data types improves forecasting accuracy and enhances model stability, particularly in environments with rapidly changing weather conditions.

CTT imagery also provides both spatial and vertical context, which supports Transformer-based models more effectively than single-point irradiance measurements. This leads to more accurate and timely forecasting of solar PV output. In addition, the strong performance of the CTT-based model in multi-step forecasting demonstrates its ability to preserve and utilize temporal patterns, making it well suited for operational planning in real-world microgrid settings.

Additionally, Figure 15 presents a comprehensive comparison of solar PV output forecasting results using different types of input data across a 300-time-step period. The figure illustrates the performance of three forecasting configurations: using solar irradiance data alone, using CTT satellite imagery with the CTT-ViT-Transformer model, and using a combined input approach that integrates irradiance, RGB imagery, and CTT data.

FIGURE 15

The results clearly demonstrate that the combined input model (red line) achieves the highest forecasting accuracy, closely aligning with the actual ground truth values (blue line) throughout the time period. This indicates that integrating multiple data sources provides a more complete and robust understanding of atmospheric conditions, allowing the model to more effectively capture complex patterns that affect solar energy generation.

In contrast, the model relying solely on solar irradiance data (orange line) shows noticeable deviations from actual output, particularly during periods of rapid fluctuation in cloud coverage. This is likely due to the spatial limitations of irradiance sensors, which capture data only at a single point and lack information on cloud dynamics.

The CTT-ViT-Transformer model (green line), which utilizes thermal cloud-top information from satellite imagery, shows substantial improvement in accuracy compared to irradiance-only models. Its ability to incorporate vertical cloud structure, such as height and temperature, enhances the model’s capacity to predict sudden drops or increases in irradiance caused by cloud movement.

The analysis confirms that CTT satellite imagery alone already provides strong predictive power, but combining it with other sources significantly improves performance. These findings support the use of multi-source input data and advanced architectures like the Transformer to enable reliable, short-term solar PV forecasting, especially in remote microgrids with limited ground-based sensing infrastructure.

For the results of using GenAI to assist Transformer-based forecasting, this study found that the hyperparameters most frequently suggested by the GenAI assistant included attention dropout values of 0.20 and 0.25 in the attention layers, feed-forward dropout values of 0.10, 0.15, and 0.20 between hidden layers, batch sizes of 16, 32, and 64, encoder depths of 2, 3, and 4 layers, decoder depths of 4, 6, and 8 layers, learning rates of 0.01, 0.001 and 4, 6, or 8 attention heads. All configurations were trained for 500 epochs to ensure sufficient convergence. The final Transformer configuration selected for the proposed framework used a batch size of 32, an encoder depth of 4 layers, a decoder depth of 8 layers, 4 attention heads, a dropout rate of 0.10, and 500 training epochs. , a learning rate of 0.01, and 500 training epochs. Compared with a baseline scenario in which wide hyperparameter ranges were specified manually, requiring 17,496 training runs and approximately 376 h (about 15.5 days) of computation on a MacBook M1 Pro (36 GB of RAM, 256 GB SSD), the GenAI-guided boundary design reduced the number of runs to 1,024 and shortened the total tuning time by 86.21% to approximately 21.5 h. Importantly, the resulting RMSE, MAE, and R2 values were not inferior to those obtained with manually defined ranges. Thus, in this framework, GenAI does not replace standard validation-based model selection but rather functions as a transparent and auditable tool for designing structured hyperparameter search spaces with substantially improved efficiency.

Furthermore, this study presents an ablation analysis that clearly isolates the roles of WGAN, CTT + ViT, and GenAI, as shown in Table 6. Enabling WGAN yields improvements over the baseline Transformer, with RMSE reduced by 2.20%, MAE reduced by 3.18%, and R2 increased by 0.0043. Replacing RGB imagery with CTT features encoded by a ViT delivers a substantial gain, reducing RMSE by 11.92%, lowering MAE by 28.63%, and increasing R2 by 0.0484, underscoring the value of physics-based thermal information. Combining WGAN with CTT + ViT achieves the highest accuracy, with RMSE reduced by 16.03%, MAE reduced by 34.54%, and R2 increased by 0.0559. Adding GenAI maintains these accuracy levels, with its primary benefit being improved training efficiency (reduced time compared with grid search) rather than further error reduction.

TABLE 6

Model variantWGANCTT + ViTGenAIMAE (kW)RMSE (kW)R2
Baseline (transformer)24.221528.87430.9227
Baseline + WGAN23.451828.23900.9270
Baseline + (CTT + ViT)17.288025.43240.9711
Baseline + WGAN + (CTT + ViT)15.854424.24450.9786
Baseline + WGAN + (CTT + ViT) + GenAI15.854424.24450.9786

Ablation results for WGAN, CTT + ViT, and GenAI.

Building on the preceding evaluation, model robustness was assessed under diverse cloud characteristics to examine performance consistency across atmospheric regimes. Forecast accuracy was stratified by two indicators: (i) the scene-average cloud-top temperature, used as a proxy for cloud-top height, and (ii) the spatial standard deviation of the CTT-derived grayscale, used as a proxy for optical and textural heterogeneity. Both indicators were min–max scaled to the range [0, 1], and the test set was partitioned into Low (0–0.3), Medium (0.3–0.7), and High (0.7–1.0) bins.

For the cloud-top height proxy (Table 7), the Low scene-average cloud-top temperature group, consistent with high, thick, and relatively stable cloud cover, yielded the lowest error (RMSE = 22.83 kW) and the highest coefficient of determination (R2 = 0.9821). Performance declined in the Medium and High groups, reflecting increased short-term irradiance variability associated with lower clouds and partly clear conditions.

TABLE 7

Mean cloud-top temperature (proxy for cloud-top height)Rangen (samples)MAE (kW)RMSE (kW)R2
Low (high, cold clouds)0–0.374914.573522.830.9821
Medium (mid-level clouds)0.3–0.736517.726425.890.9661
High (low clouds or clearer sky)0.7–1.028418.243826.810.9377

Accuracy by cloud-top height proxy (scene-average cloud-top temperature).

Extending the analysis to spatial heterogeneity (Table 8), accuracy improved with greater variability, with the High spatial SD of CTT grayscale group achieving RMSE = 22.61 kW and R2 = 0.9852. This pattern indicates that salient texture and gradient cues in fragmented cloud scenes are effectively exploited by the attention-based architecture, whereas more uniform skies provide fewer visual cues and therefore exhibit larger errors. Overall, the stratified results show that vertical cloud structure and horizontal heterogeneity jointly govern forecast difficulty, with high, cold clouds favoring stability and pronounced spatial texture further enhancing predictability. These findings confirm the robustness and applicability of the ViT–Transformer framework for sensor-limited microgrid forecasting.

TABLE 8

Grayscale cloud optical-texture variabilityRangen (samples)MAE (kW)RMSE (kW)R2
Low (smooth sky)0–0.328517.279025.280.9425
Medium (moderate variability)0.3–0.735118.193226.770.9619
High (broken clouds)0.7–1.081514.227822.610.9852

Accuracy by optical and texture variability proxy (spatial SD of CTT grayscale).

For the four-step-ahead forecasting experiment, as presented in Table 9, the model consistently achieved high accuracy across all four forecasting steps. Although there was a slight decrease in accuracy as the forecasting horizon extended, all R2 values remained above 0.96, reflecting the model’s strong predictive performance. As illustrated in Figure 16, the results further demonstrate the model’s reliability for short-term operational planning. These forward forecasts will be used to support MPC-based energy management in the operational microgrid on Koh Paluay.

TABLE 9

StepMAE (kW)RMSE (kW)R2
115.854424.24450.9786
215.928624.29220.9746
317.082225.52240.9721
418.191726.12420.9637

Forecasting results of solar PV output for 4-step ahead prediction.

FIGURE 16

Regarding generalisability and limitations, this study focuses on microgrids with solar PV generation and has been validated on the Koh Paluay microgrid, which is located in Thailand’s tropical climate. A second experiment at the AIT test site in Pathum Thani similarly confirmed a consistent relationship between cloud-top temperature (CTT) and solar irradiance. Model robustness was evaluated by stratifying scenes according to cloud-top height and cloud texture, and the results indicated stable performance across all groups.

The proposed framework can be applied to other sites with similar characteristics, regardless of microgrid size or the generation mix, although forecasting accuracy may vary with local climate and terrain. Furthermore, although the WGAN-based imputation method handles data gaps more effectively than conventional techniques, it is still recommended to have at least 3 months of historical solar PV power or irradiance data for model training and calibration.

Finally, the method primarily replaces local irradiance sensors with CTT imagery, and where ground-based sensors are available, their data can be fused with the imagery to further improve accuracy. A key limitation is the reliance on the availability and quality of CTT imagery; if some image frames are missing or corrupted, a standard Transformer model may be used as a lower-accuracy fallback. Extreme events such as tropical cyclones or tornadoes were not present in the dataset used in this study and therefore remain subjects for future investigation.

5 Conclusion

This paper introduces the CTT-ViT-Transformer, a novel forecasting framework that integrates CTT satellite imagery, Vision Transformers, and Transformer architectures to enhance short-term solar PV forecasting in microgrids without ground-based irradiance sensors. This represents the first application of CTT imagery for solar PV output forecasting and was validated with real-world data from Koh Paluay, a remote Thai island pursuing carbon neutrality.

The model outperformed conventional deep learning approaches such as LSTM and CNN-LSTM, and the use of CTT imagery produced superior results compared to RGB imagery by providing clearer cloud height information. This confirms CTT’s potential as a viable substitute for ground sensors, enabling accurate forecasting even in sensor-limited regions.

Generative AI integration enhanced data quality, corrected missing values, and optimized hyperparameters, reducing model training time while preserving prediction accuracy. The framework successfully learns both spatial and temporal patterns, delivering consistent performance across multiple time steps and proving both reliable and scalable. It can be applied to remote locations regardless of sensor availability, including areas with or without irradiance sensors or sky cameras.

Incorporating this framework into AI-driven energy management systems could reduce human intervention, improve real-time decision-making, and enhance operational flexibility. Its ability to produce accurate multi-step forecasts enables proactive energy allocation, resilience against unexpected events, and effective carbon-neutral microgrid strategies.

Overall, the CTT-ViT-Transformer offers a practical and scalable innovation for smart grids and microgrids, advancing decarbonisation, sustainable energy development, and equitable access to forecasting technologies in resource-limited regions worldwide. Future research should include cross-site and cross-climate validation, evaluation with additional satellite platforms, assessment under extreme weather events, and tighter integration with MPC-based operational control.

Statements

Data availability statement

The raw data supporting the conclusions of this article will be made available by the authors upon reasonable request.

Author contributions

SS: Conceptualization, Methodology, Data curation, Formal analysis, Writing – original draft. WO: Supervision, Methodology, Writing – review and editing. JS: Methodology, Writing – review and editing. KM: Validation, Writing – review and editing.

Funding

The author(s) declared that financial support was not received for this work and/or its publication.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fenrg.2025.1691881/full#supplementary-material

Glossary

  • ANN

    Artificial Neural Network

  • Adaboost

    Adaptive Boosting

  • BDI

    Belief–Desire–Intention

  • BP

    Back Propagation

  • CatBoost

    Categorical Boosting

  • CNN

    Convolutional Neural Network

  • CTT

    Cloud-Top Temperature

  • DHI

    Diffuse Horizontal Irradiance

  • DL

    Deep Learning

  • DNI

    Direct Normal Irradiance

  • DNN

    Deep Neural Network

  • GA

    Genetic Algorithm

  • GAN

    Generative Adversarial Network

  • GenAI

    Generative Artificial Intelligence

  • GHI

    Global Horizontal Irradiance

  • GOES-R

    Geostationary Operational Environmental Satellite – R series

  • GRNN

    General Regression Neural Network

  • GRU

    Gated Recurrent Unit

  • GTI

    Global Tilted Irradiance

  • LLM

    Large Language Model

  • LSTM

    Long Short-Term Memory

  • MAE

    Mean Absolute Error

  • ML

    Machine Learning

  • MSE

    Mean Squared Error

  • nMAE

    Normalized Mean Absolute Error

  • NLP

    Natural Language Processing

  • PCA

    Principal Component Analysis

  • PEA

    Provincial Electricity Authority

  • PV

    Photovoltaic

  • r

    Pearson correlation coefficient (r)

  • R2

    Coefficient of Determination

  • rRMSE

    Relative Root Mean Squared Error

  • RF

    Random Forest

  • RGB

    Red, Green, Blue

  • RMSE

    Root Mean Squared Error

  • RNN

    Recurrent Neural Network

  • SD

    Standard Deviation

  • SDG

    Sustainable Development Goal

  • SE

    Squeeze-and-Excitation

  • SFLA

    Shuffled Frog-Leaping Algorithm

  • SVM

    Support Vector Machine

  • TD

    Temperature Difference

  • ViT

    Vision Transformer

  • WGAN

    Wasserstein Generative Adversarial Network

  • RF

    Random Forest

References

  • 1

    AbnarS.ZuidemaW. (2020). Quantifying attention flow in transformers. arXiv Preprint arXiv:2005.00928. 10.48550/arXiv.2005.00928

  • 2

    AbumohsenM.OwdaA. Y.OwdaM.AbumihsanA. (2024). Hybrid machine learning model combining of CNN-LSTM-RF for time series forecasting of solar power generation. E-Prime – Adv. Electr. Eng. Electron. Energy9, 100636. 10.1016/j.prime.2024.100636

  • 3

    Al-AliE. M.HajjiY.SaidY.HleiliM.AlanziA. M.LaatarA. H.et al (2023). Solar energy production forecasting based on a hybrid CNN-LSTM-Transformer model. Mathematics11 (3), 119. 10.3390/math11030676

  • 4

    AlmarshoudA. (2024). Validation of satellite-derived solar irradiance datasets: a case study in Saudi Arabia. Future Sustain.2 (2), 17. 10.55670/fpll.fusus.2.2.1

  • 5

    BanikR.BiswasA. (2023). Improving solar PV prediction performance with RF-CatBoost ensemble: a robust and complementary approach. Renew. Energy Focus46, 207221. 10.1016/j.ref.2023.06.009

  • 6

    BarhmiK.HeynenC.GolroodbariS.van SarkW. (2024). A review of solar forecasting techniques and the role of artificial intelligence. Solar4 (1), 99135. 10.3390/solar4010005

  • 7

    BayasgalanO.AkisawaA. (2025). Nowcasting solar irradiance components using a vision transformer and multimodal data from all-sky images and meteorological observations. Energies18 (9), 2300. 10.3390/en18092300

  • 8

    BöckingL.MichaelisA.SchäfermeierB.BaierA.KühlN.KörnerM. F.et al (2024). Generative artificial intelligence in the energy sector.

  • 9

    BrightJ. M. (2019). Solcast: validation of a satellite-derived solar irradiance dataset. Sol. Energy189, 435449. 10.1016/j.solener.2019.07.087

  • 10

    CassolaF.BurlandoM. (2012). Wind speed and wind energy forecast through kalman filtering of numerical weather prediction model output. Appl. Energy99, 154166. 10.1016/j.apenergy.2012.03.054

  • 11

    CheferH.GurS.WolfL. (2021). “Transformer interpretability beyond attention visualization,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 782791. 10.48550/arXiv.2012.09838

  • 12

    ChiccoD.WarrensM. J.JurmanG. (2021). The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation. PeerJ Comput. Sci.7, e623. 10.7717/peerj-cs.623

  • 13

    ChuY.WangY.YangD.ChenS.LiM. (2024). A review of distributed solar forecasting with remote sensing and deep learning. Renew. Sustain. Energy Rev.198, 114391. 10.1016/j.rser.2024.114391

  • 14

    CiattoG.AguzziG.BattistiniR.BaiardiM.BurattiniS.RicciA. (2025). Exploiting GenAI for plan generation in BDI agents. Front. Artif. Intell. Appl.413, 34953502.

  • 15

    DahmaniA.AmmiY.HaniniS. (2023). A novel non-linear model based on bootstrapped aggregated support vector machine for the prediction of hourly global solar radiation. Smart Grids Sustain. Energy9 (1), 3. 10.1007/s40866-023-00179-w

  • 16

    Delgado-BonalA.MarshakA.YangY.OreopoulosL. (2022). Cloud height daytime variability from DSCOVR/EPIC and GOES-R/ABI observations. Front. Remote Sens.3, 780243. 10.3389/frsen.2022.780243

  • 17

    DeoR. C.WenX.QiF. (2016). A wavelet-coupled support vector machine model for forecasting global incident solar radiation using limited meteorological dataset. Appl. Energy168, 568593. 10.1016/j.apenergy.2016.01.130

  • 18

    DewiC.ArshedM. A.ChristantoH. J.RehmanH. A.MuneerA.MumtazS. (2024). Enhancing weather scene identification using vision transformer. World Electr. Veh. J.15 (8), 373. 10.3390/wevj15080373

  • 19

    DewiT.RismaP.OktarinaY.DwijayantiS.MardiyatiE. N.SianiparA. B.et al (2025). Smart integrated aquaponics system: hybrid solar-hydro energy with deep learning forecasting for optimized energy management in aquaculture and hydroponics. Energy Sustain. Dev.85, 101683. 10.1016/j.esd.2025.101683

  • 20

    DheerajK. D.NarayananV. L.GopalR.SharmaO.BhattaraiS.DwivedyS. K. (2025). Exploring deep learning methods for solar photovoltaic power output forecasting: a review. Renew. Energy Focus53 (March 2024), 100682. 10.1016/j.ref.2025.100682

  • 21

    DosovitskiyA.BeyerL.KolesnikovA.WeissenbornD.ZhaiX.UnterthinerT.et al (2021). “An image is worth 16×16 words: transformers for image recognition at scale,” in International conference on learning representations. 10.48550/arXiv.2010.11929

  • 22

    EjiyiC. J.CaiD.JohnsonN.Osei-MensahE.EzeF.AsareS. K.et al (2025). SolarSynthNet (SSN): a deep learning framework for binary and multiclass classification of damaged or obstructed solar panels using images. Renew. Energy256, 124224. 10.1016/j.renene.2025.124224

  • 23

    El MghouchiY. (2022). Best combinations of inputs for ANN-Based solar radiation forecasting in Morocco. Technol. Econ. Smart Grids Sustain. Energy7 (1), 27. 10.1007/s40866-022-00152-z

  • 24

    EtxegaraiG.LópezA.AginakoN.RodríguezF. (2022). An analysis of different deep learning neural networks for intra-hour solar irradiation forecasting to compute solar photovoltaic generators’ energy production. Energy Sustain. Dev.68, 117. 10.1016/j.esd.2022.02.002

  • 25

    FengH.YuC. (2023). A novel hybrid model for short-term prediction of PV power based on KS-CEEMDAN-SE-LSTM. Renew. Energy Focus47, 100497. 10.1016/j.ref.2023.100497

  • 26

    GaoH.LiuM. (2022). “Short-term solar irradiance prediction from sky images with a clear sky model,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, 24752483. 10.1109/WACV51458.2022.00313

  • 27

    GholamiA.AmeriM.ZandiM.GhoachaniR. G.KazemH. A. (2022). Predicting solar photovoltaic electrical output under variable environmental conditions: modified semi-empirical correlations for dust. Energy Sustain. Dev.71, 389405. 10.1016/j.esd.2022.10.012

  • 28

    GoodfellowI. J.Pouget-AbadieJ.MirzaM.XuB.Warde-FarleyD.OzairS.et al (2014). Generative adversarial networks. arXiv Preprint arXiv:1406.2661. 10.48550/arXiv.1406.2661

  • 29

    GuptaP.SinghR. (2021). PV power forecasting based on data-driven models: a review. Int. J. Sustain. Eng.14 (6), 17331755. 10.1080/19397038.2021.1986590

  • 30

    GuptaA. K.SinghR. K. (2022). Short-term day-ahead photovoltaic output forecasting using PCA-SFLA-GRNN algorithm. Front. Energy Res.10, 1029449. 10.3389/fenrg.2022.1029449

  • 31

    GuptaP.SinghR. (2023). Forecasting hourly day-ahead solar photovoltaic power generation by assembling a new adaptive multivariate data analysis with a long short-term memory network. Sustain. Energy, Grids Netw.35, 101133. 10.1016/j.segan.2023.101133

  • 32

    HammerA.KühnertJ.WeinreichK.LorenzE. (2015). Short-term forecasting of surface solar irradiance based on Meteosat-SEVIRI data using a nighttime cloud index. Remote Sens.7 (7), 90709090. 10.3390/rs70709070

  • 33

    HanifM. F.MiJ. (2024). Harnessing AI for solar energy: emergence of transformer models. Appl. Energy369, 123541. 10.1016/j.apenergy.2024.123541

  • 34

    HossainE.MahmoodA.da SilvaS. A. O. (2020). Seasonal solar irradiance forecasting using artificial intelligence techniques with uncertainty analysis. Sustain. Energy Technol. Assessments40, 100761. 10.1016/j.seta.2020.100761

  • 35

    HuZ.GaoY.JiS.MaeM.ImaizumiT. (2024). Improved multistep ahead photovoltaic power prediction model based on LSTM and self-attention with weather forecast data. Appl. Energy359, 122709. 10.1016/j.apenergy.2024.122709

  • 36

    HuseinM.ChungI. Y. (2019). Day-ahead solar irradiance forecasting for microgrids using a long short-term memory recurrent neural network: a deep learning approach. Energies12 (10), 1856. 10.3390/en12101856

  • 37

    HuseinM.GagoE. J.HasanB.PegalajarM. C. (2024). Towards energy efficiency: a comprehensive review of deep learning-based photovoltaic power forecasting strategies. Heliyon10 (13), e33419. 10.1016/j.heliyon.2024.e33419

  • 38

    IslamM. K.AkantoJ. M.ZeyadM.AhmedS. M. M. (2021). “Optimization of microgrid system for community electrification by using HOMER pro,” in 2021 IEEE region 10 humanitarian technology conference (R10-HTC), 15. 10.1109/R10-HTC53172.2021.9641615

  • 39

    JainS.WallaceB. C. (2019). “Attention is not explanation,” in Proceedings of NAACL-HLT 2019, 35433556. 10.48550/arXiv.1902.10186

  • 40

    KalyanK. S. (2024). A survey of GPT-3 family large language models including ChatGPT and GPT-4. Nat. Lang. Process. J.6, 100048. 10.48550/arXiv.2310.12321

  • 41

    KamangirH.KrellE.CollinsW.KingS. A.TissotP. (2024). FogNet-v2.0: explainable physics-informed vision transformer for coastal fog forecasting. Earth Space Sci. Open Archive. 10.22541/essoar.172191653.32706065/v1

  • 42

    KonstantinouM.KourtisA.NikolaidisA. (2021). Solar photovoltaic forecasting of power output using LSTM networks. Energy AI5, 100081. 10.3390/atmos12010124

  • 43

    LiD.ZhaoY.WangZ.JungC.ZhangZ. (2024). Large language model-driven structured output: a comprehensive benchmark and spatial data generation framework. ISPRS Int. J. Geo-Inf.13 (11), 405.

  • 44

    LiuZ.XuanL.GongD.XieX.ZhouD. (2025a). A long short-term memory–wasserstein generative adversarial Network-based data imputation method for photovoltaic power output prediction. Energies18 (2), 399. 10.3390/en18020399

  • 45

    LiuZ.XuanL.GongD.XieX.LiangZ.ZhouD. (2025b). A WGAN-GP approach for data imputation in photovoltaic power prediction. Energies18 (5), 1042. 10.3390/en18051042

  • 46

    Liu J.J.ZangH.ChengL.DingT.WeiZ.SunG. (2023). A Transformer-based multimodal-learning framework using sky images for ultra-short-term solar irradiance forecasting. Appl. Energy342, 121160. 10.1016/j.apenergy.2023.121160

  • 47

    LiuY.DuanS.HeX.WangH. (2023). Short-term PV power prediction based on the 24 traditional Chinese solar terms and adaboost-GA-BP model. Front. Energy Res.11, 1229695. 10.3389/fenrg.2023.1229695

  • 48

    MabasaB.LyskoM. D.MoloiS. J. (2022). Comparison of satellite-based and Ångström–prescott estimated global horizontal irradiance under different cloud cover conditions in South African locations. Solar2 (3), 354374. 10.3390/solar2030021

  • 49

    MaidinM. H. (2024). “Challenges in addressing energy injustice in ASEAN,” in Energy justice: affordable, reliable, sustainable and modern energy for all (Singapore: Springer Nature Singapore), 167180. 10.1007/978-981-97-6059-6_12

  • 50

    MansourS. H.AzzamS. M.HasanienH. M.Tostado-VelizM.AlkuhayliA.JuradoF. (2024). Wasserstein generative adversarial networks-based photovoltaic uncertainty in a smart home energy management system including battery storage devices. Energy306, 132412. 10.1016/j.energy.2024.132412

  • 51

    MercierT. M.RahmanT.SabetA. (2023). “Solar irradiance anticipative transformer,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 20652074. 10.1109/CVPRW59228.2023.00200

  • 52

    MercierT. M.SabetA.RahmanT. (2024). Vision transformer models to measure solar irradiance using sky images in temperate climates. Appl. Energy362, 122967. 10.1016/j.apenergy.2024.122967

  • 53

    MischosS.DalagdiE.VrakasD. (2023). Intelligent energy management systems: a review. Artif. Intell. Rev.56, 1163511674. 10.1007/s10462-023-10441-3

  • 54

    MohammadiE.AlizadehM.AsgarimoghaddamM.WangX.SimoesM. G. (2022). A review on application of artificial intelligence techniques in microgrids. IEEE J. Emerg. Sel. Top. Industrial Electron.3 (4), 878890. 10.1109/JESTIE.2022.3198504

  • 55

    NahidF. A.OngsakulW.MadhuM. N.LaopaiboonT. (2020). “Hybrid neural networks for renewable energy forecasting,” in Research advancements in smart technology, optimization, and renewable energy, 200222. 10.4018/978-1-7998-3970-5.ch011

  • 56

    NahidF. A.OngsakulW.ManjiparambilN. M. (2023). Short term multi-steps wind speed forecasting for carbon neutral microgrid by decomposition based hybrid model. Energy Sustain. Dev.73, 87100. 10.1016/j.esd.2023.01.016

  • 57

    OuyangZ.LiZ.ChenX. (2025). Day-ahead photovoltaic power forecasting with multi-source temporal-feature convolutional networks. Energy Inf.8 (1), 68. 10.1186/s42162-025-00531-7

  • 58

    PospíchalJ.KubovčíkM.Dirgová LuptákováI. (2022). Solar irradiance forecasting with transformer model. Appl. Sci.12 (17), 8852. 10.3390/app12178852

  • 59

    Provincial Electricity Authority (2024). Developing energy for Thai society toward a sustainable future with clean energy on koh paluay Island. Available online at: https://www.pea.co.th/news/corporate-news/628.

  • 60

    ShamsM. H.NiazH.HashemiB.LiuJ. J.SianoP.Anvari-MoghaddamA. (2021). Artificial intelligence-based prediction and analysis of the oversupply of wind and solar energy in power systems. Energy Convers. Manag.250, 114892. 10.1016/j.enconman.2021.114892

  • 61

    SharmaP.MishraR. K.BholaP.SharmaS.SharmaG.BansalR. C. (2024). Enhancing and optimising solar power forecasting in dhar district of India using machine learning. Smart Grids Sustain. Energy9 (1), 16. 10.1007/s40866-024-00198-1

  • 62

    SonY.YoonY.ChoJ.ChoiS. (2022). Cloud cover forecast based on correlation analysis on satellite images for short-term photovoltaic power forecasting. Sustainability14 (8), 4427. 10.3390/su14084427

  • 63

    SongZ.HuangL.DongQ.ZhangG.ChewM. Y. L.SetungeS.et al (2025). Impacts of shadow conditions on solar PV array performance: a full-scale experimental and empirical study. Energy320, 135219. 10.1016/j.energy.2025.135219

  • 64

    SunY.LeeJ.KimS.SeonJ.LeeS.KyeongC.et al (2023). Energy theft detection model based on VAE-GAN for imbalanced dataset. Energies16 (3), 1109. 10.3390/en16031109

  • 65

    SurathunmanunS.OngsakulW.SinghJ. G. (2024). “Exploring the role of generative artificial intelligence in the energy sector: a comprehensive literature review,” in 2024 international conference on sustainable energy: energy transition and net-zero climate future (ICUE) (IEEE), 111. 10.1109/ICUE63019.2024.10795598

  • 66

    TaoK.ZhaoJ.WangN.TaoY.TianY. (2024). Short-term photovoltaic power forecasting using parameter-optimized variational mode decomposition and attention-based neural network. Energy Sources, Part A Recovery, Util. Environ. Eff.46 (1), 38073824. 10.1080/15567036.2024.2323158

  • 67

    TawnR.BrowellJ. (2022). A review of very short-term wind and solar power forecasting. Renew. Sustain. Energy Rev.153, 111758. 10.1016/j.rser.2021.111758

  • 68

    TianF.FanX.WangR.QinH.FanY. (2022). A power forecasting method for ultra‐short‐term photovoltaic power generation using transformer model. Math. Problems Eng.2022 (1), 94214009421415. 10.1155/2022/9421400

  • 69

    TsaiW. C.TuC. S.HongC. M.LinW. M. (2023). A review of state-of-the-art and short-term forecasting models for solar PV power generation. Energies16 (14), 5436. 10.3390/en16145436

  • 70

    TzoumanikasP.NikitidouE.BaisA. F.KazantzidisA. (2016). The effect of clouds on surface solar irradiance. Renew. Energy95, 314322. 10.1016/j.renene.2016.04.026

  • 71

    VaswaniA.ShazeerN.ParmarN.UszkoreitJ.JonesL.GomezA. N.et al (2017). Attention is all you need. arXiv Preprint arXiv:1706.03762. 10.48550/arXiv.1706.03762

  • 72

    WangJ.HuW.XuanL.HeF.ZhongC.GuoG. (2024). TransPVP: a Transformer-based method for ultra-short-term photovoltaic power forecasting. Energies17 (17), 4426. 10.3390/en17174426

  • 73

    XuR.LiT.FourieC. (2025). Seasonal forecasting of solar irradiance using partial functional regression and LSTM. Solar Energy262, 3145.

  • 74

    YaoJ.ZhangL.HuangJ. (2025). Evaluation of large language model-driven AutoML in data and model management from human-centered perspective. Front. Artif. Intell.8, 1590105. 10.3389/frai.2025.1590105

  • 75

    YirgaG.FekrieD.MenberuT.WorkinehY. (2024). Time series trends and correlations of aerosol optical depth and cloud parameters over Addis Ababa. Front. Earth Sci.12, 1452075. 10.3389/feart.2024.1452075

  • 76

    ZhanG.YeY.GuoH.CaoX.TangJ.BuL. (2024). “Short-term photovoltaic power forecasting based on patch-Transformer model,” in 2024 china automation congress (CAC) (IEEE), 32813285. 10.1109/CAC63892.2024.10864684

  • 77

    ZhangW.LuoY.ZhangY.SrinivasanD. (2020). SolarGAN: multivariate solar data imputation using generative adversarial network. IEEE Trans. Sustain. Energy12 (1), 743746. 10.1109/TSTE.2020.3004751

  • 78

    ZhangB.ZhouX.YiT.ShenH.YaoY. (2023). CloudViT: a lightweight vision transformer network for remote sensing cloud detection. IEEE Geoscience Remote Sens. Lett.20 (4), 15. 10.1109/LGRS.2022.3233122

Summary

Keywords

solar forecasting, photovoltaic power prediction, microgrid, cloud top temperature, transformer, vision transformer, sensorless forecasting

Citation

Surathunmanun S, Ongsakul W, Singh JG and Mehran K (2026) Short-term solar PV forecasting in microgrids using cloud top temperature and vision transformer based models. Front. Energy Res. 13:1691881. doi: 10.3389/fenrg.2025.1691881

Received

24 August 2025

Revised

19 November 2025

Accepted

27 November 2025

Published

05 January 2026

Volume

13 - 2025

Edited by

Kok-Keong Chong, Tunku Abdul Rahman University, Malaysia

Reviewed by

Guowei Dai, Sichuan University, China

Ramiro Barbosa, Polytechnic Institute of Porto, Portugal

Updates

Copyright

*Correspondence: Weerakorn Ongsakul, ; Surasak Surathunmanun,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics