METHODS article

Front. Built Environ., 20 May 2026

Sec. Urban Science

Volume 12 - 2026 | https://doi.org/10.3389/fbuil.2026.1825552

Research on AI-assisted generation of jiangnan classical garden architectural scenes based on LLM and LDM

  • 1. School of Architecture and Urban Planning, Suzhou University of Science and Technology, Suzhou, Jiangsu, China

  • 2. School of Landscape Architecture, Jiaxing University, Jiaxing, Zhejiang, China

Abstract

Classical garden visualization commonly suffers from high production cost, time-consuming production processes, and low efficiency in iterative design. To address these limitations, this study proposes a domain-oriented AI-assisted design framework for Jiangnan classical garden architectural scenes based on a large language model (LLM) and a latent diffusion model (LDM). Rather than emphasizing the generic combination of LLMs and diffusion models, the study focuses on constructing a workflow that links natural-language parsing, structured semantic representation, distributed LoRA-based coordinated generation, and multi-level evaluation within a strongly domain-constrained design context. Specifically, the LLM is fine-tuned through supervised fine-tuning using classical garden literature and standardized samples to generate structured prompts aligned with professional design semantics, while a distributed hierarchical LoRA strategy is applied to the LDM to separate architectural morphology control from environment-level scene expression. A trigger-word-based coordination mechanism further supports rule-driven invocation and dynamic weighting of architectural and environmental LoRA models. Through this framework, the study seeks to address the unstable architectural type recognition, insufficient spatial-semantic mapping, and stylistic deviation often observed when general-purpose generation models are applied to traditional garden architectural scenes. Evaluation results indicate that more than 75% of the generated images reach a professionally usable level in spatial organization and artistic-conception restoration, while over 80% exhibit traditional garden aesthetics and visual authenticity. The proposed framework improves generation efficiency and provides a domain-oriented technical pathway for controllable generation, design iteration, design communication, and literature-informed reconstruction of Jiangnan classical garden scenes.

1 Introduction

Jiangnan classical gardens, distinguished by their unique aesthetic values and profound historical significance, represent an important embodiment of China’s outstanding traditional culture. In the conservation, restoration, and design practice of Jiangnan classical gardens, scheme refinement and visualization rely heavily on an accurate understanding of spatial relationships. Unlike modern architectural design, which typically centers on isolated building objects, architectural scenes in Jiangnan classical gardens function as spatial interfaces and visual frameworks that actively participate in the construction of the overall landscape. As an integral component of garden space, their value lies not only in the form and craftsmanship of individual elements, but more critically in the organization of multiple elements and the relationships among scenes (Qi, 2012; Liu, 2006; Xiao, 2022; Wang, 2019; Zhao, 2021).

Architectural scene visualization serves as an important medium for design exploration, scheme communication, and restoration-oriented interpretation in classical garden practice. However, existing production workflows are commonly constrained by high cost, complex production procedures, and low iteration efficiency. Early-stage visualization typically requires extensive repetitive work, including reference collection, modeling, texturing, rendering, and parameter tuning, and often relies on the coordinated use of multiple professional software tools. Moreover, conventional concept images often suffer from insufficient representational precision and cognitive discrepancies, leading to communication misunderstandings, repeated revisions, and further increases in time and trial-and-error costs. Improving the efficiency and controllability of architectural scene visualization while preserving professional accuracy therefore remains a central challenge in classical garden design.

In recent years, generative artificial intelligence (AIGC) has been increasingly applied to visual expression in architecture and landscape-related fields. Compared with earlier approaches based on generative adversarial networks (GANs), which often show limitations in generation stability and fine-grained control (Liu et al., 2024; Dhariwal and Nichol, 2021; Zhou and Liu, 2021; Chen and Zhao, 2023), latent diffusion models (LDMs) have demonstrated clearer advantages in image detail, diversity, and controllability. As a result, diffusion-based generation has become an important technical pathway in architectural and landscape conceptual design and visualization research (Chen et al., 2024; Wen et al., 2024; Li and Li, 2024; Xu, 2024; Wu et al., 2025). At the same time, large language models (LLMs), with their capacity for natural-language understanding and semantic organization, offer a useful means of reducing the operational complexity of diffusion-based image generation by transforming design descriptions into machine-readable prompt structures (Witteveen and Andrews, 2022; Ma et al., 2024).

However, the mere combination of LLMs and diffusion models does not by itself resolve the difficulties of domain-specific architectural generation. Existing LLM–diffusion workflows are largely developed for general-purpose image synthesis or broad architectural visualization tasks, and their direct application to Jiangnan classical garden architectural scenes remains problematic. In such scenes, the design object is strongly constrained by architectural typology, roof form, spatial composition, accompanying landscape elements, and regional stylistic conventions. General-purpose workflows often show unstable recognition of architectural categories, insufficient mapping from design-oriented spatial semantics to generation conditions, and stylistic deviation from the formal characteristics of traditional garden architecture. In other words, the key challenge is not simply how to connect an LLM with a diffusion model, but how to construct a domain-oriented controllable generation framework for a design object with strong semantic, morphological, and stylistic constraints.

This study addresses that challenge by developing an AI-assisted design framework for Jiangnan classical garden architectural scenes. The contribution of the study does not lie in the generic integration of an LLM and a diffusion model, but in the construction of a domain-oriented methodological framework linking natural-language parsing, structured semantic representation, distributed hierarchical LoRA generation, and multi-level evaluation. Specifically, the study makes four contributions. First, it establishes a domain-specific semantic parsing and structured prompt generation scheme grounded in classical garden design knowledge. Second, it introduces a distributed hierarchical LoRA strategy that separates architectural morphology control from environment-level scene expression. Third, it proposes a rule-based coordination mechanism linking trigger-word matching, hierarchical LoRA invocation, and dynamic weight assignment. Fourth, it constructs a multi-level evaluation framework combining expert assessment, user evaluation, VLM-assisted scoring, and error-pattern analysis. Rather than claiming novelty in the general combination of LLMs and diffusion models, this study aims to provide a domain-oriented technical pathway for controllable generation, design communication, and literature-informed reconstruction of Jiangnan classical garden architectural scenes.

2 Methods

This study adopts a deep learning–based approach to support early-stage design ideation and design communication, and constructs an end-to-end generation workflow that takes natural language as input and outputs high-quality architectural scene images. The overall framework consists of two core modules: a prompt generator and a visualization generator. The former is built upon the large language model GPT-4o (OpenAI, 14 May 2024) to parse design semantics and generate structured prompts, while the latter relies on the latent diffusion model SDXL 1.0 (Stability AI, July 2024) combined with LoRA sub-models to achieve controllable generation of architectural scenes. By integrating natural language understanding and image generation within a unified framework, the proposed workflow reduces semantic deviations caused by multi-software switching and enables professional and controllable text-to-image generation (Figure 1).

FIGURE 1

The selection of GPT-4o and SDXL 1.0 was guided by considerations of model stability, compatibility with structured semantic processing and controllable generation, and reproducibility within a standardized workflow. GPT-4o was chosen for its strong capability in natural-language understanding and structured semantic organization, which is critical for mapping design-oriented descriptions into standardized prompt schemas (OpenAI, 2024a; 2024b). SDXL 1.0 was selected because it provides well-established support for LoRA-based fine-tuning (Podell et al., 2023), which is essential for implementing the hierarchical modeling strategy adopted in this study. Although more recent diffusion models have demonstrated improved generative performance, their integration with the specific domain-oriented fine-tuning and control strategy adopted in this study was not examined here. Therefore, SDXL 1.0 was adopted as a technically stable and methodologically compatible baseline.

2.1 Prompt generator construction

The prompt generator in this study was developed to support architectural morphology recognition, scene-element extraction, and SDXL-compatible structured prompt generation. Although GPT-4o provides strong general semantic understanding, task-specific adaptation was still required for accurate interpretation of domain-specific vocabulary and architectural morphology in Jiangnan classical garden design. Therefore, prompt engineering and supervised fine-tuning (SFT; OpenAI, 2025) were jointly employed to map natural-language design descriptions to structured prompts for downstream image generation.

Two large language models were used in different stages of the prompt-generation workflow. GPT-4o (OpenAI, 14 May 2024) served as the primary experimental model for structured semantic mapping and prompt generation. Its input consisted of natural-language design requirements describing architectural type, scene atmosphere, viewpoint, and accompanying landscape elements, and its output was a structured prompt compatible with the SDXL-based image generation pipeline. In addition, GPT-4 (OpenAI, 14 March 2023) was used in a separate review-oriented role to examine whether the generated prompts conformed to predefined semantic rules, field completeness requirements, and trigger-word conventions. In this stage, the input was the structured prompt produced by GPT-4o together with the corresponding prompt schema or checking criteria, and the output was a consistency-oriented review result rather than a new generative prompt. Thus, GPT-4 functioned as a rule-based reviewer within the experimental workflow, whereas GPT-4o functioned as the main semantic generation model. The structured prompt schema and corresponding annotation rules used in this workflow are summarized in Table 1. The outputs of these two models were validated through rule-based consistency checking, manual inspection, and downstream image-generation performance.

TABLE 1

DimensionDefinitionTypical values
Trigger wordControl token associated with a specific LoRA sub-modelsingle_arch_fang; single_arch_ting_pointed.roof
Architectural typologyPrimary architectural category of the scenePavilion; hall; tower; boat pavilion; water pavilion
Morphological featuresStructural and formal characteristics of the architecturePointed roof; saddle roof; multi-eave; hexagonal
Spatial relationshipPositional relationship between architecture and surroundingsWaterside; courtyard-centered; forest interior; near rockery
Scene elementsLandscape elements co-occurring with the architectureBamboo; rockery; water; trees; stone path
Stylistic attributesCultural and aesthetic characteristics of the sceneTraditional oriental architecture; jiangnan garden style
Viewpoint and compositionViewing angle and compositional framingFront view; corner view; part view; close-up; mid-range
Technical parametersTechnical descriptors for diffusion-based generationStandard lens; wide-angle lens; soft light; overcast

Structured prompt schema for the prompt generator.

2.1.1 Fine-tuning strategy of the prompt generator

In this study, prompt engineering techniques are employed to decompose design-method inputs into key semantic units that can be reliably interpreted by the large language model. These units span multiple dimensions, including architectural morphology, structural configuration, scale, roof and decorative details, accompanying landscape elements, and scenic composition strategies. Through this field-based semantic decomposition, the model is enabled to accurately translate natural garden-design language into professional design representations without reliance on extensive domain-specific pretraining (Kampelopoulos et al., 2025), thereby providing a solid professional foundation for subsequent supervised fine-tuning.

To further enable the model to learn the semantic rules, expressive conventions, and SDXL-compatible prompt formats specific to classical garden architectural scene design (Stability AI, 2025) while preserving its general-purpose capabilities, supervised fine-tuning is conducted using the OpenAI fine-tuning interface. This process optimizes only the application layer of the model, without large-scale retraining of underlying parameters, and is therefore characterized by lightweight deployment, strong controllability, and high compliance (Dong et al., 2024). The objective of fine-tuning is to enhance the model’s ability to map natural-language design descriptions into structured prompt schemas compatible with SDXL, ensuring consistency between the outputs of the LLM and the trigger-word system of the LoRA sub-models.

2.1.2 Corpus construction and semantic annotation

To align the prompt generator with the design semantics and normative principles of classical garden architecture, the training data were constructed based on established theories and practices in classical garden design. Key design knowledge, including architectural typologies, morphological features, spatial relationships, and scene-composition principles, was systematically extracted from authoritative sources such as Yuan Ye, History of Chinese Classical Gardens, and Analysis of Chinese Classical Gardens, and then organized into structured semantic components for prompt construction. The extraction framework and representative examples are summarized in Supplementary Table S1. Based on these materials, supervised fine-tuning samples were manually compiled as paired data, each consisting of a natural-language design description and a corresponding structured prompt. The design descriptions cover key dimensions such as architectural typology, stylistic features, and accompanying landscape elements, whereas the corresponding structured prompts follow the semantic and syntactic conventions of diffusion-model inputs and incorporate trigger words aligned with the LoRA sub-model system.

2.1.3 Training, validation, and consistency review

Model training followed standard prompt engineering and supervised fine-tuning procedures. The supervised fine-tuning dataset consisted of 100 expert-curated prompt–response pairs, split into training and validation sets at a ratio of 80:20. Training data were formatted as JSONL with role-based message fields, where the user message contained a natural-language design description and the assistant message contained a structured SDXL-compatible prompt. Fine-tuning was performed through the OpenAI fine-tuning interface with a learning-rate multiplier of 2, 5 epochs, a batch size of 1, a total of 28,200 training tokens, and a fixed random seed (668018508) for reproducibility.

The validation set was used to monitor the structural stability and semantic adequacy of generated prompts, with convergence of the loss function serving as the primary stopping criterion. In addition to loss-based validation, a rule-based consistency review mechanism was introduced to assess whether the fine-tuned prompt generator could stably produce outputs conforming to the predefined prompt schema. The review was applied to prompts generated from the validation set and focused on four aspects: field completeness, lexical conformity, formatting consistency, and trigger-word consistency. Specifically, the review examined whether required semantic fields were present, whether the generated expressions conformed to the predefined controlled vocabulary, whether the prompts followed the required field order and structural format, and whether the generated trigger words matched the intended architectural categories and LoRA invocation rules.

This consistency review mechanism was designed to complement loss-based validation by examining output reliability at the structural level. Rather than evaluating visual quality directly, it was intended to verify whether the prompt generator could produce standardized and operational prompts for downstream diffusion-based generation under diverse design inputs. The review criteria are summarized in Table 2.

TABLE 2

Review dimensionDefinition
Field completenessWhether all required semantic fields are present in the generated prompt
Lexical conformityWhether generated expressions conform to the predefined controlled vocabulary
Formatting consistencyWhether prompts follow the required field order and structural format
Trigger-word consistencyWhether trigger words match the intended architectural category and LoRA invocation rules

Rule-based consistency review criteria for validation-set prompts.

2.2 Image generator construction

The image generator developed in this study focuses on synthesizing classical garden architectural morphology, stylistic characteristics, decorative details, and accompanying landscape elements.

Within the image generator, the latent diffusion model (LDM) employs a CLIP-based text–image encoder to align prompt text with corresponding visual features. Image synthesis is subsequently completed through a K-sampler and a variational autoencoder (VAE) decoder. The SDXL 1.0 model is selected as the base diffusion model. Compared with earlier diffusion models, SDXL 1.0 demonstrates substantial improvements in image sharpness, color fidelity, and fine-grained detail representation, while maintaining stable performance in high-resolution image generation tasks (Podell et al., 2023). In addition, its strong compatibility with LoRA provides sufficient flexibility for the customized fine-tuning strategy adopted in this study.

2.2.1 Model fine-tuning with LoRA

To accommodate the morphological complexity, fine-grained details, and stringent stylistic constraints of classical garden architecture, we fine-tune SDXL 1.0 using Low-Rank Adaptation (LoRA). This strategy retains the base model’s general generative capacity while efficiently learning target architectural morphologies and scene attributes via low-rank parameter updates, thereby improving controllability and output stability without substantially increasing model size (Hu et al., 2021). Moreover, LoRA enables parallel training of multiple sub-models, which supports a hierarchical modeling scheme covering both distinct architectural types and the overall environmental context. Accordingly, fine-tuning performance is strongly influenced by a well-defined architectural taxonomy (i.e., an appropriate categorization of building types) and the structural completeness and consistency of the training samples. To preliminarily assess the sensitivity of LoRA hyperparameters to architectural feature learning, controlled comparative experiments were conducted under fixed prompt, seed, and inference settings.

2.2.2 Image data collection and preprocessing

Jiangnan classical gardens take architecture as the spatial framework of the overall landscape, where architectural elements are characterized by small scale, diverse typologies, and refined decorative details. To meet the requirements of high-quality training samples for the image generator, image data were collected from two sources: real-world photographs and digital twin images.

Real-world images were captured using high-resolution cameras from multiple viewpoints and focal lengths, while digital twin images were incorporated to compensate for limitations in on-site photography. By integrating real-world and digital twin images into a unified training dataset, the proposed approach enhances feature completeness and training accuracy, while reducing the risks of sample contamination, poor convergence, and overfitting caused by noisy real-world data.

To ensure accurate generation of architectural morphology and stylistic characteristics in classical gardens, architectural samples were categorized according to Yingzao Fayuan and related literature. Specifically, buildings were classified into six primary types—halls, pavilions, towers, waterside structures, boats, and kiosks—and further subdivided into 20 detailed categories based on their spatial location within the garden, roof form, and structural configuration, thereby ensuring precise representation of architectural elements. Representative examples for each category were selected from authoritative literature, including Illustrated Practices of Yingzao Fayuan, History of Chinese Classical Gardens, Analysis of Chinese Classical Gardens, and Suzhou Classical Gardens. The classification criteria and representative examples for these architectural categories are provided in Supplementary Table S2.

2.2.3 Model training

Classical garden scenes contain diverse elements, fine-grained details, and complex spatial relationships, which make mixed multi-element fine-tuning prone to underfitting, feature interference, and training instability. A central challenge, therefore, is to preserve sensitivity to architectural details while maintaining coherent scene-level representation.

To address this issue, a distributed LoRA training strategy with two hierarchical levels was adopted, consisting of element-level and environment-level modeling. For single-building–dominant scenarios, close-up real-world images in which the primary architectural element occupied more than 60% of the image area, together with corresponding digital twin images, were used to train 20 independent LoRA sub-models, each targeting a specific architectural category. For composite scenes containing multiple architectural and landscape elements, mid- and long-range images in which the primary architectural element occupied less than 60% of the image area were aggregated to train a unified environment LoRA, so as to capture broader spatial relationships and overall stylistic consistency.

All LoRA models were trained under a unified training script using cloud-deployed LoRA scripts (Akegarasu, 2025) on the AutoDL platform, with an 8:2 training–validation split and regularization strategies applied to reduce overfitting and feature interference. During training, loss curves and validation-image quality were monitored after each epoch. Checkpoints with the lowest and most stable validation loss were manually selected to reduce the risks of underfitting, overfitting, and training instability. Final model selection was determined by jointly considering loss trends and expert visual inspection of validation images.

2.3 Performance evaluation of the image generator

To comprehensively evaluate the applicability of the fine-tuned SDXL 1.0 model for classical garden architectural visualization, we integrated the 20 element-level LoRA sub-models and one environment-level LoRA into the SDXL 1.0 inference pipeline. Using this end-to-end setup, we batch-generated 200 representative architectural scene images, each accompanied by its corresponding design description, and assessed the outputs from three complementary perspectives: expert review, user experience evaluation, and VLM-as-a-Judge–based automated scoring.

2.3.1 Human evaluation framework

Human evaluation was conducted using a structured multi-criteria scoring framework. Four evaluation dimensions were considered: accuracy of architectural form representation, rationality of spatial organization, consistency of semantic alignment, and consistency of scene atmosphere expression. Each dimension was scored on a 5-point scale, where 1 indicated very poor performance and 5 indicated excellent performance.

The weighting of the four dimensions was determined according to the research objectives and the key priorities of architectural scene generation. Accuracy of architectural form representation was assigned the highest weight (0.30), because whether the generated result correctly reflects the intended architectural type and roof form is the most fundamental criterion in classical garden architectural visualization. Rationality of spatial organization (0.25) and consistency of semantic alignment (0.25) were assigned equal intermediate weights, as both are essential for evaluating whether the generated scene is compositionally coherent and responsive to the design description. Consistency of scene atmosphere expression was assigned a slightly lower weight (0.20), because although atmosphere is important to classical garden representation, it is comparatively more subjective and more easily influenced by stylistic preference and perceptual variation among evaluators. The final evaluation result was calculated as a weighted composite score based on these four dimensions.

2.3.2 Expert and user evaluation

For expert evaluation, 10 domain experts in landscape architecture were invited, including 3 professors, 2 associate professors, 3 lecturers, and 2 senior engineers from industry. All experts had relevant academic or professional experience in traditional architectural design and landscape practice. The evaluation was conducted independently to reduce potential bias.

For user evaluation, perceptual feedback was collected through questionnaires and interviews from 23 participants with relevant backgrounds, including 20 graduate students and 3 practicing designers. Compared with expert evaluation, the user evaluation focused more strongly on perceptual readability, visual coherence, and overall scene acceptability, while still following the same general evaluative dimensions. All questionnaire and interview procedures involving human participants were conducted under ethics approval from the Ethics Committee of Suzhou University of Science and Technology (IRB 190703), and informed consent was obtained from all participants.

Using this framework, 200 representative architectural scene images generated by the end-to-end workflow were evaluated from the perspectives of architectural accuracy, spatial organization, semantic consistency, and atmosphere expression. The expert and user evaluations together provided a human-centered assessment of the model’s output quality and practical usability.

2.3.3 VLM-as-a-judge

To complement human evaluation and enable scalable assessment, two multimodal scoring models, denoted as VLM^exp and VLM^usr, were trained to simulate the tendencies of expert and user judgments, respectively. The base model was Qwen2-VL-7B-Instruct. Training was performed on an AutoDL cloud instance equipped with an NVIDIA RTX 4090 (24 GB).

The scoring model takes as input an image, its corresponding prompt or design description, and a questionnaire item, and uses five-point Likert ratings as labels. Data were split into train/validation/test = 0.8/0.1/0.1, stratified by question ID. After lightweight fine-tuning and threshold calibration, the two VLM-based evaluators were used to score a larger set of 2,000 images generated by the end-to-end workflow. To facilitate reproducibility, the full training parameter configuration of the VLM-based evaluators is provided in Supplementary Table S3.

To verify the reliability of the automated evaluation, 20 images were randomly sampled for manual cross-checking. The scoring trends produced by the VLM models were found to be broadly consistent with the corresponding expert and user questionnaire results. Given the limited size of the manual cross-check sample, the VLM-based evaluation in this study was used as a supplementary tool for trend validation and large-scale comparative assessment, rather than as a formal statistical agreement test or a replacement for human evaluation.

2.4 Module coordination and end-to-end workflow principle

An end-to-end collaborative generation mechanism was established to coordinate the prompt generator and the image generator for Jiangnan classical garden architectural scenes. The workflow consists of three stages: semantic parsing, generation control, and image output. First, natural-language design requirements are transformed into structured prompts containing architectural type, roof and decorative features, accompanying landscape elements, viewpoint, and technical parameters, thereby providing constrained semantic input for downstream generation. Second, the structured prompt is used to guide hierarchical LoRA-based generation control, so that architectural form and scene environment can be coordinated within the same inference process. Finally, the generated results are output as high-resolution architectural scene images, completing the end-to-end workflow from natural-language input to visual representation.

2.4.1 Hierarchical LoRA invocation and weight control

To ensure coordinated generation of architectural morphology and overall scene environment, a rule-based hierarchical LoRA invocation mechanism was introduced at the inference stage. In this mechanism, trigger words embedded in the structured prompt served as the primary basis for selecting the corresponding architectural LoRA sub-model. Each trigger word was uniquely associated with a specific architectural category and functioned as the key identifier for activating the element-level LoRA responsible for that building form.

For each generation task, one element-level architectural LoRA was selected as the primary model according to the trigger word specifying the target architectural type, while the environment LoRA was simultaneously loaded as an auxiliary model. The architectural LoRA mainly controlled building morphology and category-specific formal features, whereas the environment LoRA mainly contributed scene-level constraints related to vegetation, terrain, waterside composition, atmosphere, and overall stylistic coherence.

The two LoRA models were combined through a rule-driven dynamic weighting strategy rather than by fixed uniform weights. In general, the architectural LoRA was assigned a relatively higher weight, typically around 0.8, because architectural form accuracy was treated as the primary control objective in this study. By contrast, the environment LoRA was assigned a moderate auxiliary weight, usually within the range of 0.4–0.6, so as to enhance scene coherence without weakening the structural expression of the target building.

The specific weight setting was further adjusted according to the semantic characteristics of the prompt. When the prompt placed stronger emphasis on architectural form, roof type, or structural features, a higher architectural LoRA weight and a relatively lower environment LoRA weight were adopted. When the prompt included richer environmental constraints, such as waterside setting, dense vegetation, rockery composition, or atmospheric requirements, the environment LoRA weight was moderately increased to strengthen scene-level expression. Representative rules for dynamic weight assignment are summarized in Table 3.

TABLE 3

Prompt emphasisTypical architectural LoRA weightTypical environment LoRA weightIntended effect
Architectural type and roof form emphasized0.8–0.90.4–0.5Preserve architectural morphology and roof accuracy
Balanced architectural and scene constraints0.80.5Maintain both structural clarity and scene coherence
Strong environmental or atmosphere constraints0.7–0.80.5–0.6Strengthen waterside setting, vegetation, rockery, and atmosphere expression

Representative rules for dynamic weight assignment in hierarchical LoRA invocation.

3 Model training and implementation

3.1 Construction of the prompt generator

3.1.1 Sample preparation and preprocessing

The training corpus was primarily derived from chapters related to classical garden architectural scenes in authoritative treatises on traditional Chinese garden design. The collected texts mainly cover key design aspects such as architectural typology, scale, decorative details, surrounding elements, and scenic composition techniques. All textual data were carefully screened and manually corrected where necessary to ensure semantic completeness and professional accuracy.

To improve the semantic consistency and domain validity of the prompt fine-tuning dataset, a structured prompt schema was established based on a field-constrained semantic framework. Rather than treating prompts as free-form textual descriptions, each sample was organized as a sequence of semantic units corresponding to predefined schema dimensions, including trigger words, architectural typology, morphological features, spatial relationships, scene elements, stylistic attributes, viewpoint, and technical parameters, as defined in Table 1. These dimensions were aligned with both classical garden design knowledge and the structured prompt requirements of SDXL. To reduce semantic ambiguity, each dimension was associated with a controlled vocabulary, and all prompts followed a standardized ordering of semantic units during annotation.

The annotation scheme was further designed as a model-aligned semantic interface linking the prompt generator and the LoRA-based image generator. In particular, trigger words were constrained to correspond to specific LoRA sub-models, while the remaining semantic fields provided complementary constraints on morphology, spatial organization, and visual style. This design allowed the generated prompts to remain structurally standardized while being directly compatible with the downstream diffusion-based generation pipeline.

The annotation dimensions and controlled vocabulary were derived from authoritative literature on classical garden design and architectural classification. Key descriptive elements related to architectural typology, roof form, spatial organization, environmental composition, and stylistic expression were extracted from classical treatises and relevant studies, and then translated into annotation fields compatible with the structured prompt format used in this study. Based on these sources, a label system was constructed by mapping recurrent design descriptors to predefined semantic fields, while synonymous or overlapping expressions were standardized to maintain conceptual clarity.

For supervised fine-tuning (SFT), 100 paired training samples were manually constructed in accordance with the fine-tuning guidelines. Each sample consisted of a natural-language design description as input and a structured prompt compatible with SDXL as output. The samples were balanced across different scene types, architectural categories, seasonal lighting conditions, and viewpoints to improve model generalization. The dataset was split into training and validation sets at a ratio of 8:2.

To evaluate annotation reliability, a randomly selected subset of 30 samples was independently reviewed by three faculty experts from the School of Architecture, including two professors and one associate professor. Pairwise Cohen’s Kappa coefficients were calculated between the primary annotator and each expert reviewer. The resulting Kappa values ranged from 0.75 to 0.82, with a mean value of 0.78, indicating substantial agreement. Cases of disagreement were further examined and used to refine ambiguous or overlapping labels, thereby improving the consistency of the final annotation scheme.

3.1.2 Prompt engineering and supervised fine-tuning

The prompt engineering samples, stored in TXT format, were uploaded to the OpenAI development platform to construct a prompt knowledge base. Supervised fine-tuning was conducted using the OpenAI fine-tuning interface. As training progressed, both training loss and validation loss exhibited a clear downward trend and stabilized after convergence. Correspondingly, training and validation accuracy increased and remained stable after convergence (Figure 2).

FIGURE 2

3.1.3 Consistency review of the prompt generator

After fine-tuning, the model outputs on the validation set were examined using the rule-based consistency review mechanism described in Section 2.1.3. The generated prompts were checked in terms of field completeness, lexical conformity, formatting consistency, and trigger-word consistency.

The review results indicate that the fine-tuned model was able to produce prompts that consistently followed the predefined structural and formatting rules. No major violations of field composition or prompt ordering were observed in the validation outputs, and the generated trigger words remained aligned with the intended architectural categories and the corresponding LoRA invocation scheme. Minor inconsistencies were mainly related to occasional lexical variation in auxiliary scene descriptions, but these did not affect downstream prompt usability.

Overall, the consistency review suggests that the constructed prompt generator can reliably produce structured prompts for Jiangnan classical garden architectural scenes, thereby providing a stable semantic interface for subsequent image generation.

3.2 Construction of the image generator

3.2.1 Sample preparation and preprocessing

Training samples for the image generator consisted of both real-world photographs and digital twin images to balance realism and structural completeness. Real-world images were captured using a high-resolution camera, with an original resolution of 7728 × 5152 pixels. Multi-angle and multi-scale photography was employed to cover different elevations, viewpoints, and spatial distances, ensuring comprehensive representation of architectural forms and details. In addition, selected scenes were re-photographed at identical locations, angles, and focal lengths based on illustrations in classical garden literature to faithfully reproduce canonical architectural features.

Digital twin images were generated from architectural models constructed according to Illustrated Yingzao Fayuan Methods and Construction Methods and Examples of Suzhou Garden Architecture. The models were rendered using Twinmotion based on Unreal Engine 5, with materials, colors, lighting conditions, camera angles, and focal lengths aligned with real-world photographs. The dataset included 2,398 real-world images (2,226 architectural instance images and 172 re-photographed images based on literature) and 982 digital twin images, resulting in a total of 3,380 samples. The image sources therefore consisted primarily of real-world photographs, supplemented by digital twin images to improve structural completeness and category coverage.

In terms of architectural distribution, sample sizes varied across categories, reflecting differences in real-world availability, architectural complexity, and accessibility during image acquisition. More common building types, such as main halls and regular pavilion forms, accounted for a larger proportion of the dataset, whereas structurally complex or less frequently documented categories, such as multi-eaved towers and irregular pavilion types, were relatively underrepresented. Detailed sample distribution across all architectural categories is provided in Supplementary Table S4.

To support the hierarchical LoRA training strategy, the dataset was further divided into single-building–dominant samples and environment-level samples according to the proportion of the primary architectural element in the image. Close-up images in which the primary architectural element occupied more than 60% of the image area were used for category-specific element-level LoRA training, whereas mid- and long-range images in which the primary element occupied less than 60% were aggregated for training the environment LoRA. This division ensured that both architectural morphology and overall scene relationships could be learned under appropriate data conditions.

For model development, the dataset was split into training and validation subsets at a ratio of 8:2. The split was performed in a category-balanced manner to maintain the relative distribution of architectural types across the two subsets and to avoid excessive bias toward high-frequency categories. Detailed train–validation counts for each category are reported in Supplementary Table S4.

All images were preprocessed to improve data consistency and to emphasize the primary architectural elements. Using Capture One (16.5.1.14, Capture One A/S, Denmark), irrelevant regions at the edges of the images, such as adjacent buildings and tourists, were cropped to highlight the main architectural subject. Vegetation partially occluding the architecture was retained to preserve the completeness of scene elements. White balance and exposure were corrected using the AI-assisted color and exposure adjustment functions in Capture One to avoid clipping in the RGB channels and exposure histogram. Adobe Lightroom (14.1.0, Adobe Systems Incorporated, United States of America) was then used for AI-based denoising to improve image clarity. Finally, all images were uniformly resized to 1024 × 682 or 682 × 1024 pixels to ensure consistency in subsequent model training.

Each sample was manually annotated with structured prompt labels consistent with those used for prompt generator training. The annotation structure and labeling format are illustrated in Figure 3.

FIGURE 3

3.2.2 Construction of architectural scene sub-models

LoRA-based fine-tuning was performed using cloud-deployed training scripts on the AutoDL platform. Under a unified training configuration, 20 architectural LoRA sub-models and one environment LoRA model were trained in parallel. Training was constrained by both validation loss and perceptual image quality, with regularization strategies applied to mitigate overfitting and feature interference. The training and validation sets followed an 8:2 split. After each epoch, loss curves and corresponding generated examples were monitored to assess convergence behavior and visual consistency (Figure 4). Validation images were generated using SDXL 1.0 with fixed inference settings, including 25 sampling steps, a CFG scale of 7, a LoRA weight of 0.8, and a resolution of 1024 × 682 pixels. Representative training parameters are summarized in Table 4.

FIGURE 4

TABLE 4

ParameterValueParameterValue
Resolution1,024 × 682optimizer_typeAdamW8bit
network_dim64network_alpha32
learning_rate1e-4lr_schedulercosine_with_restarts
unet_lr1e-4text_encoder_lr1e-5
train_batch_size4max_train_epochs20
prior_loss_weight1mixed_precisionbf16
enable_bucketTRUESeed1337

Representative LoRA training parameters.

The selection of key LoRA hyperparameters, particularly network_dim, network_alpha, and learning_rate, was not determined arbitrarily, but was informed by pilot comparative tests on representative architectural categories. Because these parameters directly affect feature-learning capacity, adaptation strength, and convergence behavior, representative parameter combinations were compared under the same prompt, random seed, and inference settings. A controlled qualitative comparison of the generated results is presented in Figure 5.

FIGURE 5

As shown in Figure 5, lower-rank settings (network_dim = 32, network_alpha = 16) limited the model’s ability to learn fine-grained architectural features, resulting in structural inconsistencies in roof representation and insufficient expression of ridge ornaments, tile textures, and other decorative details. By contrast, the intermediate setting (network_dim = 64, network_alpha = 32) achieved a more balanced result, with smoother roof curvature, clearer hierarchical expression of the roof system, and more stable correspondence between columns and overhanging eaves. When the rank and alpha were further increased (network_dim = 128, network_alpha = 64), local details were enhanced, but the continuity of roof curvature was occasionally disrupted, and the generated outputs showed a tendency toward reduced variation, with slight signs of overfitting in local structural details.

Learning rate mainly influenced convergence efficiency and structural stability. Under a lower learning rate (5e-5), the generated results remained relatively stable, but local structural features and decorative details were insufficiently learned, indicating weak feature adaptation. Under a higher learning rate (5e-4), local roof and component details were over-enhanced, while structural consistency became less stable, as reflected in irregular tile arrangements, slight roof misalignment, and fluctuating stylistic expression across samples. In comparison, the setting of 1e-4 provided a better balance between convergence speed, structural fidelity, and stylistic consistency.

Taken together, these observations suggest that LoRA hyperparameters do not affect generation quality in a linear manner. Instead, excessively low settings tend to weaken feature learning, whereas excessively high settings may introduce instability or overfitting. Based on both the visual comparison in Figure 5 and the summarized observations in Table 5, the final configuration (network_dim = 64, network_alpha = 32, learning_rate = 1e-4) was adopted as a balanced setting for subsequent training. To facilitate reproducibility, a full training parameter configuration for a representative architectural category (Fang) is provided in Supplementary Table S5.

TABLE 5

network_dimnetwork_alphalearning_rateObserved image characteristicsTypical issuesAssessment
32161.00E-04Roof structure is not accurately learned, with insufficient representation of fine architectural detailsStructural inaccuracies; insufficient detail learningLimited feature-learning capacity
64321.00E-04Produces balanced structural fidelity with smooth roof curvature and consistent stylistic expressionMinor detail loss in complex regionsSelected setting; best balance between fidelity and stability
128641.00E-04Over-sharpened appearance with enhanced local details; slight inconsistencies observed in structural elements such as eave curvatureIncorrect fine details; misaligned eave corners; reduced variationOverfitting tendency
64325.00E-05Generates relatively stable structures, but with insufficient detail representation and weakened feature learningBlurred or simplified details; incomplete feature learningStable but underfitting
64325.00E-04Enhances local details but introduces instability in structural representation, particularly in roof alignment and stylistic consistencyStructural instability; irregular detail patterns; inconsistent styleLearning rate too high

Comparison of representative LoRA hyperparameter settings and their observed effects on generated images.

During training, underfitting was observed mainly in categories with relatively high structural complexity or greater intra-category variation, particularly hip-and-gable halls, hard-gable halls, multi-eaved towers, octagonal pavilions, and boat pavilions. These categories generally involve more complex roof hierarchies, denser decorative components, or less regular relationships among eaves, ridges, and supporting structural elements, making them more difficult to learn than simpler pavilion or hall types. In the generated validation images, underfitting was manifested as incomplete or inaccurate roof structures, blurred ridge and eave details, weakened correspondence between columns and overhanging eaves, and insufficient response to category-specific formal characteristics.

By contrast, overfitting was more likely to occur in structurally simpler categories, such as round pavilions, square pavilions, and fan-shaped pavilions. In these cases, the generated outputs tended to become visually fixed, with reduced variation across samples and repetitive formal patterns, suggesting that the models had become overly dependent on a limited range of category-specific features.

To reduce the likelihood of misattributing these problems to other factors, the annotation rules, prompt format, and training data structure were kept unchanged during tuning. The observed deficiencies were concentrated in a limited number of complex categories, while overfitting was mainly associated with several morphologically simple categories, rather than being randomly distributed across all categories. This pattern, together with the corresponding validation loss behavior and representative validation images, supports the interpretation that the main issue was related to model adaptation behavior under different category conditions, rather than annotation inconsistency or prompt ambiguity.

Targeted parameter adjustment was then performed for the affected categories. Specifically, the number of training epochs was increased, and learning rates were adjusted according to the convergence behavior of each category. For underfitting categories showing slow feature learning and weak structural response, the learning rate was moderately increased or the training duration was extended to strengthen adaptation efficiency. For categories showing unstable loss fluctuations, the learning rate was reduced and the training duration was further extended to improve convergence stability. For overfitting-prone simple categories, the training duration was controlled more conservatively to avoid excessive formal fixation. Representative tuning strategies and corresponding improvements are summarized in Table 6.

TABLE 6

Problem typeRepresentative categoriesMain manifestationsTuning strategyObserved improvement
UnderfittingHip-and-gable halls; hard-gable halls; multi-eaved towers; octagonal pavilions; boat pavilionsIncomplete or inaccurate roof structures; blurred ridge and eave details; weakened correspondence between columns and overhanging eaves; insufficient response to category-specific formsTraining duration was extended, and the learning rate was adjusted according to convergence behavior to improve feature learning and structural stabilityValidation loss became more stable; roof hierarchy became clearer; ridge and eave details were better preserved; roof curvature and structural correspondence were improved
OverfittingRound pavilions; square pavilions; fan-shaped pavilionsReduced variation across samples; repetitive formal patterns; visually fixed outputs caused by over-dependence on limited category-specific featuresTraining duration was controlled more conservatively, and the learning rate was reduced when necessary to suppress excessive formal fixationRepetitive formal patterns were alleviated; visual variation across samples increased; stylistic expression became more flexible and coherent

Targeted tuning strategies and observed improvements for underfitting- and overfitting-prone categories.

After tuning, the affected categories showed more stable validation loss trends and improved visual performance. In particular, roof hierarchy became clearer in hip-and-gable halls and multi-eaved towers, roof curvature and structural correspondence were improved in octagonal pavilions and boat pavilions, and repetitive formal patterns were alleviated in round, square, and fan-shaped pavilions.

To further examine the effectiveness of parameter tuning for the affected categories, expert assessment was conducted on representative validation outputs before and after optimization. The evaluation focused on four aspects: accuracy of architectural form representation, rationality of spatial organization, consistency of semantic alignment, and consistency of scene atmosphere expression. Scores were assigned on a 5-point scale by 5 domain experts, and the mean values are summarized in Table 7.

TABLE 7

CategoryEvaluation criterionBefore tuningAfter tuning
Hip-and-gable hallsAccuracy of architectural form representation3.13.8
Rationality of spatial organization33.6
Consistency of semantic alignment2.84.2
Consistency of scene atmosphere expression3.33.7
Multi-eaved towersAccuracy of architectural form representation34.1
Rationality of spatial organization2.93.5
Consistency of semantic alignment3.14.2
Consistency of scene atmosphere expression3.23.6
Octagonal pavilionsAccuracy of architectural form representation2.83.5
Rationality of spatial organization2.93.6
Consistency of semantic alignment2.84
Consistency of scene atmosphere expression3.13.5

Expert evaluation before and after parameter tuning for representative affected architectural categories.

As shown in Table 7, the optimized models achieved higher scores across all four evaluation dimensions. Among them, semantic alignment showed the largest improvement, indicating that parameter tuning enhanced the model’s ability to respond more consistently to category-specific design semantics. Improvements were also observed in architectural form representation and spatial organization, suggesting that the optimized settings contributed to more stable structural expression in the affected categories. By contrast, the increase in scene atmosphere expression was relatively moderate, implying that atmosphere-related perception remained more sensitive to scene complexity and stylistic variation. These results support the effectiveness of the targeted tuning strategies for the affected architectural categories.

3.3 End-to-end generation workflow

Based on the trained models, an end-to-end generation workflow integrating the prompt generator and image generator was implemented using the ComfyUI platform (Comfy-Org, 2026). Natural-language design requirements are first processed by the prompt generator to produce structured prompts. According to trigger information within the prompts, corresponding architectural LoRA sub-models and the environment LoRA are selectively loaded to generate images consistent with the design intent. Finally, the generated images are upscaled to produce high-resolution classical garden architectural visualizations.

4 Results

4.1 Overall generation performance

The proposed end-to-end assisted design system is capable of stably generating architectural scene images that conform to the design principles and formal characteristics of Jiangnan classical gardens based on natural-language design inputs (Figure 6). The results indicate that the prompt generator consistently produces structurally valid prompts and accurately activates the corresponding architectural and environment LoRA models. The generated images exhibit strong semantic consistency with the design descriptions in terms of architectural form, decorative details, spatial organization, and arrangement of surrounding elements.

FIGURE 6

Further analysis demonstrates that the proposed system maintains robust performance across varying levels of design semantic complexity (Figure 7). In different design stages, the system is able to correctly invoke appropriate architectural typologies and preserve compositional stability and coherent organization of scene elements, indicating good adaptability to diverse design intentions.

FIGURE 7

In terms of generation efficiency, the system achieves relatively high throughput while preserving professional quality and controllability. On a local workstation equipped with an RTX 3080 GPU, the average generation time is approximately 1–1.5 min per image (40–60 images per hour). On a cloud-based AutoDL platform with an NVIDIA A100 GPU, the generation time is reduced to 10–20 s per image (180–360 images per hour). These results verify the effective coordination between the prompt generator and the image generator, as well as the controllability and efficiency of the proposed workflow for Jiangnan classical garden architectural scene generation.

4.2 Reliability evaluation and optimization of the image generator

4.2.1 Comparison with baseline diffusion models

The fine-tuned SDXL model was qualitatively compared with several widely used diffusion models without domain-specific fine-tuning, including SDXL 1.0, FLUX.1 dev, and DALL·E 3, using identical prompts and generation parameters. The purpose of this comparison was to provide a reference for the behavior of general-purpose models in the absence of domain-specific adaptation, rather than to establish a comprehensive performance benchmark across different diffusion architectures.

The results indicate that, when directly applied to Jiangnan classical garden architectural scenes, general-purpose models may produce irrelevant elements, incorrect architectural forms (e.g., inappropriate pavilion or hall types), non-conforming vegetation layouts, or incomplete scene components such as rockeries or streams (Figure 8). Within the proposed SDXL-based workflow, the incorporation of LoRA fine-tuning improves the alignment between generated results and domain-specific architectural forms and scene elements, particularly in the generation of canonical structures such as octagonal pointed-roof pavilions.

FIGURE 8

4.2.2 Human evaluation results

Based on the multi-criteria human evaluation framework described in Section 2.3, expert assessment showed that the generated images achieved generally positive results across the four evaluation dimensions of architectural form representation, spatial organization, semantic alignment, and scene atmosphere expression. More than 75% of the generated images reached a professionally usable level in spatial organization and scene atmosphere expression, while over 80% performed positively in architectural form representation and semantic alignment. User surveys involving graduate students and practitioners yielded comparable trends, indicating that the generated results were generally well received in both professional and user-oriented evaluation contexts.

4.2.3 VLM-as-a-judge evaluation

The VLM-as-a-Judge evaluation produced scoring trends that were broadly consistent with the corresponding expert and user assessments. When applied to a larger set of 2,000 generated images, the automated evaluation supported the overall stability of model performance across architectural accuracy, semantic consistency, and visual coherence. Manual cross-checking on 20 randomly sampled images suggested that the VLM-based results were generally aligned with human judgment. In this study, the VLM-based evaluation was used primarily for trend validation and large-scale supplementary assessment, rather than as a formal statistical agreement test.

4.3 Error pattern analysis of generated results

To further characterize the limitations of the proposed workflow, the generated validation outputs were systematically reviewed and categorized according to recurrent error patterns. Based on manual examination of 677 validation images, four error types were identified, including roof structure error, decorative detail loss, structural misalignment, and semantic inconsistency. The frequencies of these error types are summarized in Table 8. Because a single generated image could contain more than one error type, the reported percentages do not sum to 100%.

TABLE 8

Error typeDescriptionFrequency, n (%)More common inPossible cause
Roof structure errorIncomplete or inaccurate roof hierarchy, ridge misplacement, incorrect eave geometry83 (12.3%)Hip-and-gable halls; multi-eaved towersHigh structural complexity; insufficient feature learning
Decorative detail lossMissing or blurred ridge ornaments, timber details71 (10.5%)Complex halls; octagonal pavilionsWeak local-detail representation
Structural misalignmentInconsistent relation between columns, eaves, and roof edges49 (7.2%)Boat pavilions; towersIrregular structural relationships
Semantic inconsistencyGenerated results do not fully match prompt-specified viewpoint, spatial orientation, or accompanying scene elements36 (5.3%)Prompts involving complex spatial constraints or multiple scene-element requirementsIncomplete semantic parsing and weak alignment between auxiliary prompt conditions and visual generation

Classification and frequency of typical errors in generated architectural scenes.

As shown in Table 8, roof structure error and decorative detail loss were the most common failure types, occurring in 12.3% and 10.5% of the validation outputs, respectively. Structural misalignment was observed in 7.2% of the images, whereas semantic inconsistency was relatively less frequent, accounting for 5.3%. Overall, the most recurrent problems were concentrated in architectural structure and fine-grained detail representation, rather than in global scene atmosphere or overall visual coherence.

Representative failure cases are illustrated in Figure 9. Roof structure errors were mainly observed in architecturally complex categories, such as hip-and-gable halls and multi-eaved towers, where the generated results sometimes exhibited incomplete roof profiles, inaccurate ridge placement, or inconsistent eave curvature. Decorative detail loss was more common in categories requiring dense fine-grained ornamentation, such as octagonal pavilions, and was manifested in blurred tile textures, weakened ridge ornaments, or simplified timber details. Structural misalignment was particularly evident in boat pavilions and certain tower-like forms, where the relationships among columns, roof edges, and supporting elements were not always consistently maintained. Semantic inconsistency was mainly observed in cases where the textual description involved multiple constraints or category-specific expressions, indicating that semantic parsing and image generation were not always fully aligned.

FIGURE 9

These findings suggest that the remaining limitations of the current workflow are not randomly distributed across outputs, but instead show clear associations with architectural complexity, local detail density, and category-specific structural irregularity. In particular, categories with complex roof hierarchies and dense decorative components remain more challenging for the model to learn, whereas semantic inconsistency tends to be associated with prompts involving multiple auxiliary constraints. This error analysis provides a more detailed basis for understanding the current performance boundaries of the workflow and offers guidance for further model refinement and data organization.

5 Conclusion

This study develops a domain-oriented AI-assisted design framework for Jiangnan classical garden architectural scenes by integrating a large language model (GPT-4o) with a latent diffusion model (SDXL 1.0). The contribution of the work does not lie in the generic combination of LLMs and diffusion models, but in the construction of a methodological framework that connects natural-language parsing, structured semantic representation, distributed hierarchical LoRA generation, and multi-level evaluation within a strongly domain-constrained design context. Through this framework, the study attempts to address several recurrent limitations of general-purpose generation models in traditional garden architectural scenes, including unstable architectural type recognition, insufficient spatial-semantic mapping, and stylistic deviation from regional design characteristics.

The results indicate that the proposed workflow can achieve end-to-end generation from natural-language input to architectural scene visualization, while maintaining comparatively stable performance in architectural form representation, semantic consistency, and generation efficiency. It should be noted that the present study focuses on validating the effectiveness of the proposed framework within an SDXL-based workflow, rather than establishing a comprehensive performance comparison across different diffusion architectures. In this sense, the workflow provides a feasible technical pathway for design exploration, visual communication, and literature-informed scene reconstruction in the context of Jiangnan classical gardens.

It should also be emphasized that the present study is not primarily concerned with image quality in a general visual sense, but with the ability of AI methods to respond to design semantics, architectural typology, and spatial relationships in traditional garden architectural scenes. The main question is therefore not whether the system can generate more visually attractive images, but whether the generated results can reflect, to a certain extent, the structural constraints and domain-knowledge requirements embedded in architectural design tasks. From this perspective, the study is more appropriately positioned as research on an AI-assisted design methodology for classical garden architectural scenes than as a purely visual generation experiment.

Several limitations should nevertheless be noted in relation to the reported results. First, although the overall generation performance was generally stable, the model still showed weaker performance in architecturally complex categories. As indicated by the error analysis, roof structure error and decorative detail loss remained the most frequent failure types, especially in categories such as hip-and-gable halls, multi-eaved towers, octagonal pavilions, and boat pavilions. This suggests that the current dataset and training strategy are still insufficient for fully capturing highly complex roof hierarchies, dense decorative components, and irregular structural relationships. Future work should therefore focus on expanding the coverage of structurally complex categories, improving the balance of training samples across architectural types, and introducing more targeted data augmentation or category-specific fine-tuning strategies.

Second, the current prompt-to-image workflow still has limited capacity in handling auxiliary semantic constraints beyond core architectural form. Although the LoRA-based framework can reliably control building typology, the error analysis shows that semantic inconsistency still occurred in a portion of the generated results, particularly when prompts involved viewpoint, spatial orientation, or multiple scene-element requirements. In addition, the post-optimization expert evaluation showed that improvement in scene atmosphere expression was relatively limited compared with the gains observed in structural accuracy and semantic alignment. These findings suggest that the present workflow is more effective in controlling architectural morphology than in coordinating higher-level environmental composition and atmosphere-related expression. Future research may address this limitation by refining semantic decomposition rules, improving the interaction between element-level and environment-level control, and exploring more flexible prompt-parsing and scene-constraint mechanisms.

Third, the present study remains primarily oriented toward two-dimensional visual generation and early-stage design support. The proposed workflow has shown practical value in concept generation, design communication, and literature-informed reconstruction, but it does not yet support three-dimensional structural reasoning, parametric design integration, or interactive design iteration in a comprehensive manner. In this sense, the current system should be regarded as an assisted visualization tool rather than a substitute for conventional design workflows. Future work may extend the framework toward multimodal and three-dimensional generation, integrate stronger structural constraints, and further examine its applicability in broader scenarios such as restoration design, heritage interpretation, and interactive design systems. In addition, extending the proposed framework to other diffusion architectures and systematically evaluating their performance under comparable fine-tuning and control strategies remains an important direction for future research.

Statements

Data availability statement

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Ethics statement

The studies involving humans were approved by Ethics Committee of Suzhou University of Science and Technology (IRB 190703). The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.

Author contributions

ZL: Supervision, Project administration, Methodology, Conceptualization, Validation, Investigation, Writing – review and editing, Data curation, Funding acquisition, Writing – original draft, Formal Analysis, Software, Resources, Visualization. QZ: Supervision, Writing – review and editing, Writing – original draft, Methodology, Data curation, Formal Analysis, Investigation, Validation, Conceptualization, Visualization, Project administration. LD: Conceptualization, Writing – original draft, Funding acquisition, Visualization, Validation, Writing – review and editing, Supervision, Data curation, Project administration, Methodology. MS: Validation, Methodology, Investigation, Resources, Conceptualization, Writing – review and editing, Supervision, Visualization, Formal Analysis, Software, Project administration, Writing – original draft, Funding acquisition, Data curation.

Funding

The author(s) declared that financial support was received for this work and/or its publication. This work was supported by the Jiangsu Province Graduate Student Practice and Innovation Program under the project A Classical Garden Visualization Generator Based on Large Language Models and Diffusion Models (Grant No. SJCX24_1948), and by the Intangible Cultural Heritage Consortium Project (Grant No. 4126113002).

Acknowledgments

The authors would like to thank the editors and reviewers for their valuable comments and constructive suggestions. The authors also gratefully acknowledge the support provided by their institution during the course of this research.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was used in the creation of this manuscript. During manuscript preparation, generative AI tools were used only for language-related assistance. Specifically, GPT-5.2 (OpenAI) was used for English translation, language polishing, and clarity improvement of the manuscript text. It was not used for generating experimental data, model outputs, evaluation results, or scientific conclusions. By contrast, the AI models involved in the experimental workflow are described separately in the Methods section. In this study, GPT-4o and GPT-4 were used only within the experimental framework for prompt generation and output review under predefined rules, and were not involved in post hoc writing assistance reported in this disclosure. All AI-assisted edits were reviewed and verified by the authors, who take full responsibility for the content of the final manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fbuil.2026.1825552/full#supplementary-material

References

  • 1

    Akegarasu (2025). lora-scripts. GitHub Repository. Available online at: https://github.com/Akegarasu/lora-scripts (Accessed July 15, 2025).

  • 2

    ChenR.ZhaoJ. (2023). Generation and design feature recognition of landscape architecture scheme based on style-based generative adversarial network. Landsc. Archit.30 (7), 1221. 10.12409/j.fjyl.202305050212

  • 3

    ChenR.LuoX.HeY.ZhaoJ. (2024). Research on the adaptability of generative algorithm in generative landscape design. Landsc. Archit.31 (9), 1223. 10.3724/j.fjyl.202404120207

  • 4

    Comfy-Org (2026). ComfyUI. GitHub Repository. Available online at: https://github.com/Comfy-Org/ComfyUI (Accessed March 5, 2026).

  • 5

    DhariwalP.NicholA. (2021). Diffusion models beat GANs on image synthesis. Adv. Neural Inf. Process. Syst.34, 87808794.

  • 6

    DongG.YuanH.LuK.LiC.XueM.LiuD.et al (2024). “How abilities in large language models are affected by supervised fine-tuning data composition,” in Proceedings of the 62nd annual meeting of the association for computational linguistics (ACL 2024). 10.48550/arXiv.2310.05492

  • 7

    HuE. J.ShenY.WallisP.Allen-ZhuZ.LiY.WangS.et al (2021). LoRA: low-rank adaptation of large language models. arXiv 2106.09685. 10.48550/arXiv.2106.09685

  • 8

    KampelopoulosD.TsanousaA.VrochidisS.KompatsiarisI. (2025). A review of LLMs and their applications in the architecture, engineering and construction industry. Artif. Intell. Rev.58, 250. 10.1007/s10462-025-11241-7

  • 9

    LiP.LiB. (2024). “Generating daylight-driven architectural design via diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops (CVPRW 2024). 10.48550/arXiv.2404.13353

  • 10

    LiuT. (2006). Experiencing architecture: study on Chinese scholar gardens of jiangnan in the ming and qing dynasties. Shanghai, China: Tongji University.

  • 11

    LiuY.WuP.LiX.MoW. (2024). Application and renovation evaluation of Dalian’s industrial architectural heritage based on AHP and AIGC. PLOS ONE19 (10), e0312282. 10.1371/journal.pone.0312282

  • 12

    MaB.ZongZ.SongG.LiH.LiuY. (2024). Exploring the role of large language models in prompt encoding for diffusion models. arXiv 2406.11831. 10.48550/arXiv.2406.11831

  • 13

    OpenAI (2024a). GPT-4o system card. Available online at: https://openai.com/index/gpt-4o-system-card/(Accessed March 26, 2026).

  • 14

    OpenAI (2024b). Introducing structured outputs in the API. Available online at: https://openai.com/index/introducing-structured-outputs-in-the-api/(Accessed March 26, 2026).

  • 15

    OpenAI (2025). Supervised fine-tuning (SFT) guide. Available online at: https://platform.openai.com/docs/guides/supervised-fine-tuning (Accessed July 15, 2025).

  • 16

    PodellD.EnglishZ.LaceyK.BlattmannA.DockhornT.MüllerJ.et al (2023). SDXL: improving latent diffusion models for high-resolution image synthesis. arXiv 2307.01952. 10.48550/arXiv.2307.01952

  • 17

    QiJ. (2012). Private garden space research. Xinxiang, China: Henan Normal University.

  • 18

    Stability AI (2025). Stable diffusion 3.5 prompt guide. Available online at: https://stability.ai/learning-hub/stable-diffusion-3-5-prompt-guide (Accessed July 15, 2025).

  • 19

    WangA. (2019). A comparative study of the spatial sequences of jiangnan classical gardens and modern display buildings. Master’s thesis. Jinan, China: Shandong Jianzhu University.

  • 20

    WenM.LiangD.YeH.TuH. (2024). Architectural facade design with style and structural features using stable diffusion model. J. Intelligent Constr.2 (4), 9180034. 10.26599/JIC.2024.9180034

  • 21

    WitteveenS.AndrewsM. (2022). Investigating prompt engineering in diffusion models. 10.48550/arXiv.2211.15462

  • 22

    WuY.MengL.LiL. (2025). Aided greenway design approach based on internet big data and AIGC fine-tuning model. Landsc. Archit.32 (7), 96105. 10.3724/j.fjyl.202410290624

  • 23

    XiaoL. (2022). A study on the spatial form and construction techniques of the lingering garden based on transparency theory. Master’s thesis. Jinan, China: Shandong Jianzhu University. 10.27273/d.cnki.gsajc.2022.000353

  • 24

    XuY. (2024). Exploration of small and medium scale landscape architecture spatial layout scheme generation based on stable diffusion models. Archit. Cult. (9), 268271. 10.19875/j.cnki.jzywh.2024.09.081

  • 25

    ZhaoE. (2021). Inheritance and innovation of jiangnan classical garden techniques in modern architectural environmental design: a case study of the new suzhou museum designed by I. M. Pei. Anhui Archit.28 (5), 1718. 10.16330/j.cnki.1007-7359.2021.05.008

  • 26

    ZhouH.LiuH. (2021). Artificial intelligence aided design: landscape plan recognition and rendering based on deep learning. Chin. Landsc. Archit.37 (1), 5661. 10.19775/j.cla.2021.01.0056

Summary

Keywords

generative artificial intelligence, jiangnan architecture, large language model, latent diffusion model, LORA, traditional garden design

Citation

Li Z, Zhao Q, Dong L and Sun M (2026) Research on AI-assisted generation of jiangnan classical garden architectural scenes based on LLM and LDM. Front. Built Environ. 12:1825552. doi: 10.3389/fbuil.2026.1825552

Received

08 March 2026

Revised

31 March 2026

Accepted

27 April 2026

Published

20 May 2026

Volume

12 - 2026

Edited by

Izuru Takewaki, Kyoto Arts and Crafts University, Japan

Reviewed by

MengNan Shi, Sichuan University, China

Abdulrahman Salem, Ain Shams University, Egypt

Updates

Copyright

*Correspondence: Minkai Sun,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics