Abstract
Introduction:
Educational supervision in Gulf Cooperation Council (GCC) school systems has expanded beyond classroom observation to encompass teacher evaluation, school inspection, quality assurance, instructional leadership, mentoring, coaching, professional learning, and technology-supported supervision. However, evidence remains dispersed across several related literatures.
Methods:
This systematic review synthesised 40 studies addressing direct clinical, instructional, and developmental supervision; teacher evaluation and data-driven appraisal; school inspection and quality assurance; instructional and middle leadership; and coaching, mentoring, onboarding, and professional learning. The review followed a transparent systematic procedure informed by PRISMA and qualitative evidence-synthesis principles, including explicit eligibility criteria, database searching, screening, extraction, appraisal, and thematic synthesis.
Results:
Six cross-cutting themes were identified: supervision as an institutional continuum, feedback quality, role ambiguity and distributed supervision, inspection–development tension, data and technology, and professional-learning integration. Supervision emerged as a multi-level institutional practice shaped by accountability pressures, leadership mediation, feedback processes, and professional-learning capacity.
Discussion:
The review develops a GCC-sensitive conceptualisation of supervision ranging from compliance-oriented monitoring to developmental capacity building. Effective reform requires stronger feedback literacy, protected professional dialogue, clearer differentiation between evaluation and coaching, and evidence-informed alignment among inspection, leadership development, and teacher professional learning.
1.1 1 Introduction
International and conceptual foundations
Educational supervision remains a central mechanism through which education systems seek to improve instructional quality, support teachers’ professional growth, and connect school-level practice with broader educational goals (, pp. 8–12; , pp. 27–30). In its improvement-oriented form, supervision is not limited to monitoring teachers or checking compliance; rather, it is concerned with helping teachers examine practice, interpret evidence about learning, refine instruction, and develop professionally over time (, pp. 4–8; , pp. 15–19). This makes supervision internationally significant because it sits at the intersection of teaching quality, school leadership, professional learning, and educational accountability (, pp. 31–35; UNESCO, 2024, pp. 17–22).
The international supervision literature has historically distinguished between bureaucratic or administrative supervision and developmental approaches that emphasize instructional dialogue, teacher reflection, and professional growth (, pp. 72–79; , pp. 36–41). Clinical supervision, for example, is commonly associated with structured cycles of pre-observation conference, classroom observation, data analysis, post-observation conference, and reflective follow-up (, pp. 9–15). Developmental supervision similarly assumes that teachers differ in experience, expertise, commitment, and professional needs, and therefore require differentiated forms of supervisory support (, pp. 82–91). These models remain influential because they frame supervision as an educative process rather than as a purely evaluative act (, pp. 42–45).
At the same time, contemporary education systems increasingly operate under intensified expectations for accountability, standards alignment, measurable outcomes, and school improvement (, pp. 375–378; , pp. 47–49). Global policy reports emphasize the importance of leadership, teacher support, and coherent professional learning systems in improving educational quality, but they also show that leadership and accountability arrangements vary significantly by national context (; UNESCO, 2024, pp. 23–31). UNESCO's teacher policy guidance similarly treats teacher development as a policy priority that must be embedded within coherent systems of preparation, support, career development, and professional standards (UNESCO, 2019, pp. 28–36). These global debates make educational supervision an important topic for systematic review because supervision is one of the practical mechanisms through which policy expectations are translated into teacher learning and classroom practice (Timperley, 2008, pp. 6–10; UNESCO, 2019, pp. 42–46).
1.2 The GCC context and international relevance
The Gulf Cooperation Council (GCC) is a regional intergovernmental organization established in 1981 by Bahrain, Kuwait, Oman, Qatar, Saudi Arabia, and the United Arab Emirates to advance coordination, integration, and interconnection among its member states (). The GCC provides a significant context for examining the relationship between supervision, inspection, accountability, leadership, and teacher learning because its education systems have experienced sustained reform pressures linked to national development, human capital formation, quality assurance, curriculum modernization, and school improvement (, pp. 1–8; , pp. 15–23). Reform trajectories across the six systems are not identical, and it would be methodologically inaccurate to represent the GCC as a single uniform education system (, pp. 24–31; , pp. 108–111). Nevertheless, several GCC systems have engaged with policy tools such as standards-based reform, school evaluation, leadership development, teacher professionalization, and external quality assurance (, pp. 17–28; , pp. 112–115).
These reforms have often involved the adaptation of international models to local governance, cultural, and institutional conditions (, pp. 31–38; , pp. 32–39). In such contexts, supervision cannot be understood simply as a technical process imported from international literature; it must be examined as a policy-mediated and culturally situated practice (, pp. 22–26; , pp. 115–117). For example, instructional leadership, teacher appraisal, school inspection, and professional development may be formally separated in policy, yet closely connected in school-level practice when leaders and teachers experience them as part of the same accountability environment (, pp. 381–385; , pp. 53–56).
The evidence base assembled for this review indicates that GCC supervision research is distributed across studies of instructional supervision, school inspection, teacher evaluation, quality assurance, instructional leadership, mentoring, professional development, and school reform. This distribution is important because it suggests that supervision is often present in the literature indirectly, embedded within adjacent constructs rather than always named as “clinical supervision” or “developmental supervision.” Consequently, the review treats GCC supervision as a broad but carefully bounded construct that includes direct supervisory practice as well as closely related forms of instructional monitoring, feedback, coaching, evaluation, and professional learning where they are explicitly connected to teacher improvement.
Supervision in GCC systems requires context-sensitive analysis because imported supervision models may be reshaped by centralized governance, hierarchical authority structures, inspection expectations, school leadership roles, teacher workforce composition, and national reform priorities (, pp. 39–45; , pp. 40–48). Clinical supervision assumes that teachers and supervisors can engage in evidence-informed, trust-based, reflective dialogue about classroom practice (, pp. 37–43). Developmental supervision assumes that supervisors can differentiate their approaches according to teachers’ professional readiness, expertise, and needs (, pp. 82–91). Instructional leadership assumes that school leaders are able to influence teaching and learning through shared goals, curriculum coordination, observation, feedback, and professional learning structures (, pp. 229–234; , pp. 659–663).
These assumptions may not always align neatly with reform-intensive or accountability-oriented policy environments (, pp. 383–387; , pp. 54–56). When supervision is associated with inspection readiness, performance ratings, or compliance documentation, teachers may interpret feedback less as professional support and more as evaluative surveillance (, pp. 21–27; , pp. 46–49). Conversely, when supervision is embedded within trusting professional relationships and connected to teachers’ instructional problems, it may support reflective practice and professional learning (, pp. 174–181; Timperley, 2008, pp. 18–24). The same supervisory structure may therefore produce different outcomes depending on relational trust, role clarity, evidence use, leadership capacity, and accountability consequences (, pp. 386–388; UNESCO, 2024, pp. 45–51).
A context-sensitive GCC analysis is therefore necessary for two reasons. First, it prevents overgeneralization from international supervision models developed in different professional and governance contexts (, pp. 22–26; , pp. 49–51). Second, it enables the review to examine GCC supervision as a hybrid practice shaped by developmental aspirations, inspection pressures, leadership mediation, and institutional-cultural conditions. This framing allows the review to move beyond a simple binary between “supportive supervision” and “accountability inspection” and instead examine how these logics interact in actual reform environments (, pp. 380–387; , pp. 52–56).
1.3 Conceptual boundaries of supervision
Supervision, in its instructional and developmental sense, refers primarily to the support of teacher learning, reflection, instructional improvement, and professional growth (, pp. 4–7; , pp. 10–14). Inspection, by contrast, usually refers to external or system-level evaluation of school quality against standards, indicators, or regulatory expectations (, pp. 375–377; , pp. 48–50). Teacher appraisal generally refers to the formal evaluation of individual teacher performance for accountability, employment, promotion, professional standards, or performance-management purposes (; , pp. 45–52; UNESCO, 2019, pp. 67–72).
This distinction is analytically important because the same school-level practice may carry different meanings depending on the policy context in which it occurs (, pp. 380–383). A classroom observation may function as developmental supervision when it is used for formative feedback and professional learning, but it may function as appraisal or inspection preparation when it is linked to judgement, rating, compliance, or external accountability (, pp. 21–27; , pp. 52–54). The educational value of supervision therefore depends not only on whether observation and feedback occur, but also on how they are framed, who conducts them, what evidence is used, how teachers experience them, and whether the process strengthens professional learning (, pp. 112–119; Timperley, 2008, pp. 10–15).
Teacher professional learning research further suggests that professional development is most likely to influence practice when it is sustained, evidence-informed, connected to student learning, and embedded in teachers’ actual instructional problems (Timperley, 2008, pp. 15–24). Supervision can contribute to such learning when it creates structured opportunities for observation, feedback, inquiry, coaching, reflection, and collaborative problem solving (, pp. 29–36; , pp. 174–181). However, supervision can lose its developmental function when it is dominated by performativity, fear of judgement, role ambiguity, or accountability pressure (, pp. 383–387; , pp. 54–56). For this reason, reviews of supervision must examine not only supervisory models but also the institutional logics that shape how those models are enacted (, pp. 45–49).
1.4 Review gap
As mentioned before, empirically, GCC research on supervision is fragmented across studies that examine supervision directly and studies that address related constructs such as school inspection, teacher evaluation, professional development, instructional leadership, mentoring, and quality assurance. This fragmentation makes it difficult to determine when supervision is functioning as developmental support, instructional guidance, appraisal, compliance monitoring, or inspection preparation. The problem is not only terminological; it affects the interpretation of evidence because different supervisory purposes imply different standards for judging effectiveness (, pp. 4–8; , pp. 33–36).
Theoretically, much of the available supervision theory has been developed through international models that do not always fully account for GCC policy conditions, including centralized reform, external accountability, national quality agendas, and culturally mediated authority relations (, pp. 72–79; , pp. 40–48). Clinical and developmental supervision provide valuable conceptual resources, but they require adaptation when applied to systems where supervision may be closely associated with inspection, appraisal, or performance management (, pp. 21–27; , pp. 383–387). There is therefore a need to connect supervision theory with policy enactment, instructional leadership, inspection research, teacher professional learning, and institutional analysis (, pp. 229–235; Timperley, 2008, pp. 15–24).
Regionally, GCC supervision research remains under-synthesized in the international literature. Existing international discussions of supervision, inspection, and instructional leadership often draw primarily on North American, European, Australasian, or broader comparative policy contexts, with limited attention to how supervision is translated in Gulf education systems (, pp. 27–31; UNESCO, 2024, pp. 31–37). Conversely, GCC-focused studies may provide rich contextual insights without always positioning those findings in relation to wider supervision theory. This review addresses that gap by treating GCC evidence not merely as regional description but as a basis for refining international understandings of supervision under conditions of policy borrowing, reform intensification, and accountability pressure.
1.5 Purpose and research questions
The purpose of this review is to synthesize and critically interpret the available evidence on educational supervision in GCC education systems, with particular attention to how supervision is conceptualized, enacted, and mediated by inspection, accountability, instructional leadership, appraisal, and teacher professional learning. The review does not assume that clinical, developmental, or instructional supervision models can be transferred directly into GCC systems without adaptation. Instead, it examines how such models are reframed by policy structures, school leadership arrangements, cultural expectations, and accountability environments (, pp. 37–43; , pp. 82–91; , pp. 383–387).
The review also aims to move beyond a narrow description of supervisory practices by developing a GCC-sensitive conceptual model of supervision beyond inspection. This model treats supervision as a contextually mediated improvement practice situated between policy accountability, school leadership, or coaching structures, and teacher professional learning. This framing is consistent with international evidence that improvement-oriented leadership and teacher learning depend on alignment among policy, organizational conditions, professional relationships, and instructional practice (, pp. 663–668; Timperley, 2008, pp. 18–24; UNESCO, 2024, pp. 45–51).
This review is guided by the following research questions:
How is educational supervision conceptualized and enacted across the GCC studies included in the review?
How do supervision, inspection, accountability, teacher appraisal, and professional learning interact in GCC education systems?
What evidence exists regarding clinical, developmental, instructional, and inspection-oriented forms of supervision in GCC contexts?
What institutional, cultural, policy, and leadership conditions enable or constrain developmental and improvement-oriented supervision?
How can the GCC evidence base contribute to a context-sensitive conceptual model of supervision beyond inspection?
These questions are designed to support both descriptive synthesis and theoretical interpretation. They allow the review to identify what the evidence shows, where the evidence remains limited, and how GCC supervision research can inform wider debates about supervision, accountability, leadership, and teacher learning.
1.6 Contributions and article structure
This review makes four contributions. First, it clarifies the conceptual boundaries among supervision, inspection, appraisal, accountability, and teacher professional learning in GCC education contexts. This distinction is necessary because supervision can be misinterpreted when observation, feedback, inspection preparation, and performance evaluation are treated as equivalent practices. Second, the review synthesizes a fragmented regional evidence base and calibrates claims according to evidence type, methodological strength, and contextual relevance. This is important because the GCC literature does not yet support uniform claims about the effectiveness of supervision models across all countries or school systems. Third, the review develops a GCC-sensitive conceptual argument that supervision in the region is best understood as a hybrid practice shaped by developmental aspirations, inspection pressures, leadership mediation, and institutional-cultural conditions. This contribution extends supervision theory by showing that clinical and developmental models require contextual adaptation in centralized, reform-intensive, and accountability-oriented systems. Fourth, the review positions GCC supervision research as internationally relevant by using the region as a case for examining how global supervision models are translated, constrained, and reconfigured in policy environments shaped by reform, accountability, and national development priorities.
2 Materials and methods
2.1 Review design and reporting standards
This review was designed as a systematic qualitative evidence synthesis of recent literature on educational supervision and supervision-adjacent practices in Gulf Cooperation Council (GCC) school systems. A qualitative evidence synthesis design was appropriate because the review sought to interpret concepts, mechanisms, experiences, practices, and contextual conditions across heterogeneous sources rather than calculate pooled intervention effects or aggregate standardized effect sizes (Thomas and Harden, 2008, pp. 1–2, 4–8; Tong et al., 2012, pp. 1–3). The review therefore used thematic synthesis as the principal synthesis approach because thematic synthesis is designed to move from coded findings to descriptive themes and then to higher-order analytical themes (Thomas and Harden, 2008, pp. 4–8).
The review was informed by PRISMA 2020 to support transparent reporting of the search, screening, eligibility, and selection procedures. PRISMA 2020 was appropriate because it provides updated guidance for reporting why a systematic review was conducted, what methods were used, how studies were identified and selected, and what evidence was synthesized (, pp. 1–3). In this review, PRISMA 2020 guided the documentation of information sources, search procedures, duplicate removal, title-and-abstract screening, full-text eligibility assessment, exclusion decisions, and the final movement of records into the included corpus (, pp. 4–7).
Because the review also involved interpretive synthesis of qualitative, mixed-methods, dissertation-based, policy-facing, and conceptual evidence, reporting was additionally informed by ENTREQ. ENTREQ was used because qualitative evidence synthesis requires explicit reporting of the review question, search strategy, appraisal process, data extraction, synthesis method, and transparency procedures (Tong et al., 2012, pp. 1–3). ENTREQ was not used as a rigid checklist but as a complementary transparency framework to strengthen reporting of interpretive decisions, coding procedures, theme generation, and audit-trail documentation (Tong et al., 2012, pp. 2–5).
The combined use of PRISMA 2020 and ENTREQ was methodologically appropriate because the review addressed a heterogeneous evidence base in which supervision was distributed across direct supervision studies, teacher evaluation and appraisal, school inspection and quality assurance, instructional and middle leadership, coaching, mentoring, onboarding, and professional learning (, pp. 1–3; Thomas and Harden, 2008, pp. 4–8; Tong et al., 2012, pp. 1–4). This design enabled the review to preserve systematic transparency while also allowing interpretive synthesis across differently labelled but conceptually related supervisory practices.
2.2 Search strategy and information sources
Searches were conducted and updated during the first half of 2026 till May 2026 across ProQuest, Scopus, and ERIC using database-adapted combinations of terms related to instructional supervision, educational supervision, developmental supervision, clinical supervision, teacher evaluation, teacher appraisal, supervisory feedback, lesson observation, school inspection, instructional leadership, coaching, mentoring, professional learning, and GCC school contexts. The search covered 2015–2026 and focused on school-level evidence from Bahrain, Kuwait, Oman, Qatar, Saudi Arabia, and the United Arab Emirates.
Because supervision in GCC scholarship is distributed across adjacent literatures, the strategy deliberately combined direct supervision terms with related terms for evaluation, inspection, quality assurance, instructional leadership, coaching, mentoring, and professional learning. This prevented the review from excluding relevant school-level evidence simply because authors used evaluation or inspection terminology rather than the word supervision.
Search strings were adapted to each database's indexing conventions. To strengthen transparency, the reproducible database-adapted search logic used for the final reconstruction is reported in Table 1 and in the supplementary material file.
Table 1
| Component | Final documentation |
|---|---|
| Search period | 2015-2026; searches conducted and updated during the first half of 2026 till May 2026. |
| Core information sources | ProQuest Dissertations & Theses Global, Scopus, and ERIC. Supplementary retrieval involved targeted full-text verification, citation checking, and retrieval of GCC school-level dissertations, theses, journal articles, and relevant conceptual chapters. |
| Core search concepts | Instructional supervision OR educational supervision OR clinical supervision OR developmental supervision OR teacher supervision OR supervisory feedback OR teacher evaluation OR teacher appraisal OR principal evaluation OR lesson observation OR school inspection OR school self-evaluation OR quality assurance OR instructional leadership OR coaching OR mentoring OR professional learning |
| Geographic/context terms | GCC OR Gulf Cooperation Council OR Bahrain OR Kuwait OR Oman OR Qatar OR Saudi Arabia OR United Arab Emirates OR UAE OR Abu Dhabi OR Dubai OR Sharjah |
| Population/setting terms | School OR schools OR K-12 OR teacher OR teachers OR principal OR principals OR school leader OR middle leader OR head of department OR supervisor OR inspector |
| Database counts | ProQuest: 1,400 records; Scopus: 581 records; ERIC: 346 records; total: 2,327 records. After duplicate removal, 1,324 unique records were screened. |
Search strategy, information sources, and database counts.
Database searching was supplemented by targeted full-text retrieval, citation checking, and verification of the final included-study corpus. Supplementary and conceptual sources were retained only where they contributed directly to the review questions or clarified the accountability, supervision, evaluation, or professional-learning architecture shaping GCC school systems.
Earlier foundational literature was used only for conceptual framing, methodological justification, and international comparison; it was not counted as part of the included GCC evidence corpus unless it met the school-level GCC inclusion criteria.
2.3 Eligibility criteria
The review focused on school-level educational supervision in GCC contexts. For this review, the GCC was defined as Bahrain, Kuwait, Oman, Qatar, Saudi Arabia, and the United Arab Emirates. School-level education included public and private K–12 schooling, including primary, middle, secondary, and cross-phase studies where school supervision, teacher evaluation, instructional leadership, school inspection, or professional learning was substantively relevant. The eligibility framework is provided in Supplementary Table S2.
The phenomenon of interest was defined broadly but not indiscriminately. A source was eligible when it addressed at least one of the following: clinical supervision, instructional supervision, developmental supervision, supervisory feedback, teacher–supervisor relationships, e-supervision, teacher evaluation, appraisal, data-driven teacher assessment, internal evaluation, inspection, institutional review, quality assurance, school self-evaluation, instructional leadership, middle leadership, lesson observation, coaching, mentoring, onboarding, induction, or structured professional learning. This inclusive but bounded definition was necessary because supervision in GCC school systems may be enacted through several adjacent institutional mechanisms rather than through a single formally named model of supervision.
Per that, sources were included when they met four core criteria. First, the source had to be located in at least one GCC education system or explicitly compare GCC school-system practice. Second, the source had to address school-level practice rather than higher education, general policy, or non-school professional training unless the connection to K–12 schooling was explicit. Third, the source had to examine supervision or a supervision-adjacent mechanism with sufficient conceptual proximity to the review definition. Fourth, the source had to contribute empirical, practice-based, policy-analytic, conceptual, or system-level evidence relevant to the review questions.
Sources were excluded when they discussed educational leadership, teacher professional development, reform, or policy only in general terms without a clear connection to supervisory practice, classroom observation, evaluation, feedback, accountability, teacher support, or teacher-capacity development. Sources were also excluded when they were outside the GCC region, focused exclusively on higher education without school-level relevance, lacked sufficient bibliographic or substantive information for assessment, duplicated another record, or contributed only generic commentary without analytic value for the review.
The final included corpus comprised 40 sources organized into five conceptual evidence groups: direct clinical, instructional, and developmental supervision; teacher evaluation, appraisal, and data-driven teacher assessment; school inspection, institutional review, quality assurance, and accountability; instructional leadership and middle leadership as developmental supervision; and coaching, mentoring, onboarding, and professional learning. This grouping was used to avoid conflating supervision with inspection, appraisal, leadership, or professional development while still recognizing their conceptual and institutional intersections in GCC school systems.
2.4 Screening and selection process
Screening followed a staged PRISMA-informed process. Database searches identified 2,327 records: 1,400 from ProQuest, 581 from Scopus, and 346 from ERIC. Duplicate records were removed through DOI and exact-title matching, followed by manual checking of normalised titles and duplicate full-text versions. This process removed 1,003 duplicates and left 1,324 unique records for title-and-abstract screening.
At title-and-abstract stage, 1,193 records were excluded because they did not meet the review's population, context, phenomenon, publication-type, date-range, or school-level relevance criteria. A total of 130 full texts were assessed for eligibility, and 40 sources were included in the final synthesis.
The screening process deliberately avoided narrow keyword dependency. A study of inspection, for example, was not automatically treated as a supervision study; it became relevant only when it showed how system-level review influenced school-level monitoring, feedback, professional learning, leadership, or improvement. Similarly, professional-learning studies were included only when they related to coaching, mentoring, appraisal, supervision, instructional leadership, or school-based improvement mechanisms.
2.5 Treatment of full text review decisions
The authors conducted a thorough full-text assessment of all potentially eligible records and subsequently examined each of the 40 included sources in full. For every included source, the authors reviewed the study context, school level, research purpose, design, participants or document type, methods, findings, methodological transparency, publication or repository status, and relevance to the research questions. Sources were excluded if they did not provide substantive evidence on supervision or closely related practices in GCC K–12 settings, served only as general background, or could not be reliably identified or retrieved.
Following the detailed full-text review, each of the 40 included sources was assigned an explicit inclusion rationale and an evidence-directness designation. Direct evidence comprised studies in which supervision, supervisory feedback, clinical or developmental supervision, instructional supervision, teacher–supervisor relations, or e-supervision was the primary focus. Indirect evidence included studies in which supervision was examined through teacher or principal evaluation, inspection, quality assurance, school self-evaluation, accountability, lesson observation, data-informed appraisal, or instructional leadership. Developmental evidence comprised studies of coaching, mentoring, induction, onboarding, teacher-led appraisal, lesson study, and professional learning.
This source-by-source examination ensured that each of the 40 studies was included on a clear and documented basis and that direct supervision research was distinguished from related evaluation, accountability, and professional-learning evidence. The procedure also strengthened the transparency and traceability of the synthesis in line with ENT was distinguished from related evaluation, accountability, and professional-learning evidenceREQ guidance on reporting appraisal, extraction, and analytical decisions (Tong et al., 2012, pp. 2–5).
2.6 Data extraction
Data was extracted using a structured evidence extraction matrix. For each included source, the matrix captured author, year, country or emirate where applicable, education phase, school sector, study design, participants or document type, supervision-related construct, supervisory mechanism, key findings, implications for developmental supervision, methodological limitations, evidentiary classification, and inclusion rationale. The full included-study register and condensed extraction matrix are provided in Supplementary Tables S5, S6.
For empirical studies, extracted data included research purpose, research questions, design, sample, participant groups, data collection methods, analytic procedures, key findings, supervision-related claims, and relevance to the review questions. For dissertations and theses, extraction additionally considered the level of methodological detail, sampling transparency, analytic procedures, and suitability for inclusion in a developing regional evidence base. For policy documents, inspection frameworks, and grey literature, extraction focused on issuing body, country or system coverage, policy purpose, supervisory or accountability structure, assigned roles, implementation assumptions, and relevance to school-level supervision. For conceptual sources, extraction focused on theoretical contribution, conceptual clarification, and relevance to interpreting supervision, feedback, accountability, or instructional leadership.
The extraction matrix recorded both the evidence-directness designation and the substantive evidence cluster assigned to each of the 40 sources. These fields served different purposes: the three evidence-directness categories indicated how closely each source addressed supervision, whereas the five evidence clusters organised studies according to their principal topic. This distinction supported consistent source weighting, thematic synthesis, and calibration of the review's claims.
This layered extraction structure was necessary because the review did not assume that all supervision-adjacent sources were equally direct. It allowed the synthesis to identify how supervision is explicitly studied, how it is indirectly enacted through evaluation and inspection, and how it is developmentally expressed through teacher learning and school leadership practices. This distinction also supported later evidence-confidence judgements by separating direct supervision evidence from contextual or developmental evidence.
2.7 Methodological quality appraisal
Empirical studies were appraised using a fit-for-purpose cross-design approach informed by the Mixed Methods Appraisal Tool. MMAT was appropriate because it was designed for the appraisal stage of systematic mixed-studies reviews that include qualitative, quantitative, and mixed-methods studies (, pp. 1–3). Its cross-design structure allowed the review to appraise studies with different designs without imposing a single inappropriate standard across qualitative, quantitative, and mixed-methods evidence (, pp. 1–4).
For qualitative studies, appraisal considered coherence between research questions, data sources, data collection procedures, analytic strategy, participant evidence, and interpretation. This reflected MMAT's attention to the appropriateness of qualitative approaches, adequacy of data collection, derivation of findings from data, substantiation of interpretation, and coherence across qualitative components (, pp. 3–5). For quantitative and descriptive studies, appraisal considered sampling clarity, measurement appropriateness, analytic transparency, and whether conclusions were proportionate to the data, consistent with MMAT criteria for quantitative descriptive and non-randomized study designs (, pp. 4–7). For mixed-methods studies, appraisal considered the adequacy of the qualitative and quantitative components, the rationale for integration, integration of findings, and coherence between design and interpretation, consistent with MMAT criteria for mixed-methods research (, pp. 7–8).
For dissertation-based studies, appraisal also considered whether the source provided sufficient methodological detail to support inclusion. This additional attention was necessary because dissertations and theses can provide rich regional evidence but vary in methodological transparency, publication status, and external peer-review exposure. Dissertation-based evidence was therefore retained when it contributed substantively to the review questions, but it was interpreted cautiously and weighted according to methodological transparency and evidentiary directness.
MMAT appraisal was used to contextualize claims rather than exclude studies automatically. This decision was consistent with the review's purpose: to map and interpret a developing regional evidence base rather than produce an effects-only synthesis restricted to high-design studies. It was also consistent with the logic of mixed-studies appraisal, where quality assessment can inform interpretation and confidence rather than function solely as a mechanical exclusion filter (, pp. 1–3).
Non-empirical sources were not appraised with MMAT because MMAT is designed for empirical studies. Instead, policy documents, inspection frameworks, grey literature, and conceptual sources were assessed using supplementary judgement criteria: relevance to the review questions, document authority, contextual specificity, transparency of purpose, conceptual proximity to supervision, and evidentiary function. Policy documents were treated as evidence of formal policy architecture rather than evidence of implementation effectiveness; conceptual sources were treated as interpretive resources rather than empirical confirmation of practice.
2.8 Evidence confidence and source weighting
Evidence confidence was assessed narratively at the theme level rather than expressed as a numerical score. This decision was appropriate because the review synthesized different evidence types and because methodological strength, conceptual relevance, country coverage, evidentiary directness, and explanatory value could not be reduced defensibly to a single quantitative index. The confidence logic therefore considered recurrence, coherence, directness, methodological transparency, country coverage, and explanatory value. Theme-level confidence judgements and source-weighting rules are detailed in Supplementary Tables S7, S10.
Recurrence referred to whether similar findings appeared across multiple studies, countries, or evidence clusters. Coherence referred to whether findings fitted together logically across supervision, evaluation, inspection, leadership, and professional learning literatures. Directness referred to whether the evidence addressed supervision explicitly or only through adjacent mechanisms. Methodological transparency referred to the clarity and credibility of the study design, data sources, analysis, and interpretation. Country coverage referred to whether a theme was supported across multiple GCC contexts or concentrated in one system. Explanatory value referred to whether a finding helped explain how supervision functions in practice.
Themes supported by multiple empirical studies across more than one GCC context and by consistent findings were assigned higher confidence. Themes supported by limited country coverage, indirect evidence, dissertation-heavy evidence, policy-contextual evidence, or mixed methodological quality were assigned moderate, limited, or contextual confidence. This calibration was necessary to prevent overclaiming and to ensure that the strength of conclusions matched the strength, directness, and consistency of the evidence.
The review did not apply a rigid hierarchy in which peer-reviewed journal articles automatically overrode dissertations or institutional studies. Instead, each source was judged according to relevance, transparency, directness, and contribution to the review questions. This decision was important because excluding all non-journal sources would have narrowed the review in a way that misrepresented the regional evidence base. At the same time, the synthesis did not treat all sources as equally strong. Claims based on dissertation-based, source-type-limited, policy, or grey-literature evidence were qualified and triangulated wherever possible.
Lower-confidence evidence was not automatically excluded because it often provided insight into underrepresented countries, emerging practices, policy conditions, or supervision-adjacent mechanisms. Instead, lower-confidence evidence influenced interpretation through cautious wording, triangulation, and theme-level confidence ratings. This approach allowed the review to remain inclusive while avoiding unsupported claims about the prevalence or effectiveness of supervision models across all GCC systems.
2.9 Data coding and thematic synthesis
Thematic synthesis proceeded through extraction, initial coding, consolidation of overlapping codes, development of descriptive categories, and generation of analytical themes, consistent with the logic of thematic synthesis. During the final revision audit, the code-to-theme trail was reconstructed from the frozen included-study corpus to make the synthesis more transparent for readers and reviewers.
The reconstructed audit identified 118 extracted meaning units relevant to the review questions. These generated 72 initial code labels. After duplicate labels, synonymous codes, and highly overlapping concepts were merged, 28 consolidated codes were retained. These were organised into 12 descriptive categories and interpreted as six cross-cutting analytical themes. The thematic-synthesis coding trail and code-ount audit are summarised in Table 2.
Table 2
| Synthesis stage | Audit result | Purpose in the review |
|---|---|---|
| Extracted meaning units | 118 | Study-level findings, propositions, mechanisms, constraints, and reported implications relevant to the research questions. |
| Initial code labels | 72 | First-cycle labels applied to supervision, evaluation, inspection, leadership, feedback, data use, and professional-learning evidence. |
| Consolidated codes | 28 | Merged labels after removing duplication, synonyms, and highly overlapping concepts across studies. |
| Descriptive categories | 12 | Intermediate groupings such as feedback-dialogue, formative/summative evaluation, role architecture, inspection translation, evidence routines, and job-embedded learning. |
| Analytical themes | 6 | Cross-cutting interpretive themes explaining how supervision functions across GCC school systems. |
Thematic-synthesis coding trail and code-count audit.
Coding combined deductive and inductive elements. Deductive codes were derived from supervision theory and included observation, conferencing, feedback, formative evaluation, summative evaluation, instructional leadership, coaching, mentoring, inspection, evidence use, professional learning, and role distribution. Inductive codes captured GCC-specific mechanisms and constraints, including inspection readiness, performativity, evaluator credibility, teacher voice, supervisory workload, expatriate-teacher transition, and technology-mediated supervision.
The final six analytical themes were: supervision as an institutional continuum; feedback quality; role ambiguity and distributed supervision; inspection-development tension; data and technology; and professional-learning integration. These themes are reported as formal synthesis findings, while the Discussion interprets their relationships through higher-order lenses concerning accountability, developmental supervision, leadership mediation, institutional and sociocultural conditions, and professional learning.
Contradictory findings were retained rather than smoothed away. For example, inspection was treated both as a possible improvement trigger and as a possible source of performativity; teacher evaluation was treated both as a route to feedback and as a source of compliance; and data systems were treated both as standardising tools and as mechanisms requiring ethical, professional, and contextual mediation.
2.10 Reflexivity, ethical representation, and audit trail
Reflexivity was addressed through explicit documentation of interpretive decisions made during searching, screening, appraisal, coding, theme development, and claims calibration. Reflexive documentation was necessary because the review included heterogeneous evidence and required conceptual judgements about whether supervision-adjacent practices such as inspection, appraisal, instructional leadership, mentoring, or professional learning should be included. ENTREQ-informed reporting supports this kind of transparency by requiring reviewers to make synthesis processes, appraisal decisions, and interpretive procedures explicit (Tong et al., 2012, pp. 2–5).
The audit trail included search records, recoverable database-level counts, reconstructed database search logic, screening decisions, duplicate-removal records, eligibility rationales, text-review notes, evidence extraction matrices, MMAT appraisal summaries, coding records, theme-development notes, evidence-confidence ratings, source-weighting decisions, sensitivity-check results, and claims-calibration notes. This audit-trail structure was used to make the review process traceable and to distinguish direct supervision evidence from indirect, developmental, policy-contextual, and conceptual evidence.
Each included source was assigned an inclusion rationale so that readers could see why it was treated as part of the supervision evidence base. This procedure was particularly important because some studies addressed multiple mechanisms. For example, a study of lesson observation could be relevant to accountability, instructional leadership, and supervision, but it was classified according to its primary contribution to the review.
Reflexive attention was also given to the risk of conceptual overextension. The review did not assume that inspection, appraisal, instructional leadership, professional development, and supervision are identical constructs. Rather, it examined how these constructs overlap, diverge, and interact in GCC school systems. This distinction was important because supervision-related practices may be institutionally intertwined even when they remain conceptually distinct in international literature.
Ethical considerations in this review concerned representation and interpretation rather than direct human-subject data collection. The review used already available studies and documents and did not collect new participant data. However, ethical care was still necessary because the synthesis interpreted GCC school systems, cultural contexts, and professional practices. The analysis therefore avoided deficit language and did not frame GCC contexts as deficient because they differ from Western supervision models. Instead, GCC systems were treated as complex reform contexts in which accountability, national development agendas, professional diversity, expatriate teaching workforces, school improvement pressures, and imported supervisory models interact.
The discussion of imported supervision models was therefore framed cautiously. Clinical, developmental, or instructional supervision models were not treated as universal templates to be transferred unchanged into GCC systems. Rather, they were used as conceptual resources for interpreting how supervision may require adaptation to centralized governance structures, inspection regimes, professional hierarchies, cultural expectations, and school-level capacity conditions.
2.11 Sensitivity and robustness checks
Sensitivity and robustness checks were conducted to examine whether the main thematic findings remained stable under different inclusion and weighting conditions. These checks assessed whether the synthesis depended disproportionately on particular source types, lower-confidence evidence, grey literature, dissertations, policy documents, indirect supervision studies, or country-specific evidence. Sensitivity testing was used to align the strength of claims with the stability and directness of the evidence, consistent with PRISMA-informed expectations for transparent reporting of synthesis decisions and limitations (, pp. 6–8). The complete sensitivity-analysis matrix is presented in Supplementary Table S9.
The sensitivity checks examined the stability of findings under several conditions: retaining only peer-reviewed empirical studies; retaining only high- and moderate-confidence empirical studies; excluding dissertations and theses; excluding grey literature and policy reports; focusing only on sources published from 2020 onward; analyzing country-specific evidence without GCC-wide generalization; retaining only sources directly focused on supervision; and excluding sources with limited methodological transparency.
Themes that remained visible across multiple sensitivity conditions were treated as more robust. Themes that weakened when dissertation-based, grey-literature, policy, or indirect evidence was removed were retained but interpreted cautiously. Themes dependent on limited country coverage were presented as emerging or contextual rather than conclusive. This process helped ensure that the final synthesis did not overstate the strength of the evidence or treat all GCC systems as uniform.
Sensitivity checks also informed the Results, Discussion, and Limitations sections. For example, where the evidence was stronger for supervision as linked to inspection, accountability, teacher evaluation, and instructional leadership than for fully developed clinical or developmental supervision models, the review calibrated its claims accordingly. Lower-confidence evidence was not excluded automatically, but its influence on the synthesis was made explicit through source weighting and cautious interpretation.
2.12 Limitations, safeguards, and transferability
This review has its limitations, each of which was addressed through methodological safeguards. First, the evidence base was conceptually heterogeneous. The included sources examined supervision through related but distinct constructs, including instructional supervision, developmental supervision, clinical supervision, teacher evaluation, principal evaluation, inspection, lesson observation, instructional leadership, coaching, mentoring, professional learning, school self-evaluation, and data-driven appraisal. This breadth was necessary because supervision in GCC school systems is distributed across several institutional practices rather than located in one role or model only. To manage this heterogeneity, the review classified sources by evidence directness. Direct supervision studies were given greater interpretive weight when analysing supervisory practice, while teacher-evaluation, inspection, professional-learning, data-use, and conceptual sources were used to explain the wider institutional conditions that shape supervision. Empirical studies were appraised using MMAT-informed criteria, while conceptual and policy-facing sources were treated as contextual or interpretive support rather than equivalent empirical evidence.
Second, the available literature was uneven across GCC countries, school sectors, and supervision models. UAE, Saudi Arabia, Oman, and Kuwait were more strongly represented than Bahrain and Qatar, and several studies focused on particular public, private, or international-school contexts. To address this, the review avoids presenting every finding as uniformly applicable to all GCC systems. Claims are framed as recurrent patterns within the included corpus, with country-specific and sector-specific boundaries noted where relevant. The findings should therefore be understood as analytically transferable to comparable GCC and accountability-oriented school contexts.
Third, much of the available evidence was based on interviews, surveys, document analysis, perception data, or cross-sectional designs. These studies are valuable for understanding how teachers, principals, supervisors, and school leaders experience supervision, evaluation, feedback, inspection, and professional learning. However, they provide limited basis for causal claims about teacher change or student outcomes. The review therefore does not claim that any single supervisory model automatically improves teaching or achievement across GCC systems. Instead, it identifies the mechanisms and conditions under which supervision-related practices appear more likely to become developmental: feedback quality, evaluator credibility, role clarity, teacher voice, leadership capacity, contextualised evidence use, and sustained professional learning.
Taken together, these safeguards strengthen the trustworthiness of the synthesis. The review offers a calibrated account of how supervision-related practices operate across the included GCC evidence base and explains the conditions under which accountability-oriented systems may become more developmental. Its contribution lies in identifying patterns, mechanisms, and contextual conditions.
3 Results
3.1 Evidence-base profile and study selection
The final corpus comprised 40 studies organised into five substantive evidence clusters. These clusters differ from the three evidence-directness categories used during full-text review and extraction: the clusters describe each study's principal topic, whereas the direct, indirect, and developmental designations indicate its relationship to supervision. Nine sources addressed clinical, instructional, developmental, supervisory, or e-supervision; eight examined teacher evaluation, appraisal, or data-informed teacher assessment; nine investigated school inspection, institutional review, quality assurance, school self-evaluation, or accountability; eight focused on instructional or middle leadership as a form of developmental supervision; and six addressed coaching, mentoring, onboarding, personalised professional learning, or teacher professionalism. This distribution shows that GCC supervision research spans several related institutional literatures rather than forming a single, self-contained field. Figure 1 presents the study-selection flow, and Table 3 summarises the five evidence clusters.
Figure 1
Table 3
| Evidence cluster | Studies (n) | Synthesis function |
|---|---|---|
| Direct clinical, instructional, and developmental supervision | 9 | Defines supervision through observation, supervisory feedback, teacher–supervisor relations, clinical/developmental supervision, and e-supervision. |
| Teacher evaluation, appraisal, and data-driven assessment | 8 | Shows how supervision is formalised through appraisal standards, internal evaluators, teacher-led appraisal, and data systems. |
| School inspection, institutional review, quality assurance, and accountability | 9 | Represents the system-level supervisory architecture of external review, school self-evaluation, quality assurance, and accountability. |
| Instructional and middle leadership as developmental supervision | 8 | Shows school-based supervision enacted through principals, heads of department, senior teachers, lesson-study facilitators, and distributed leaders. |
| Coaching, mentoring, onboarding, and professional learning | 6 | Shows supervision as professional capacity building through coaching, mentoring, induction, personalised learning, and teacher agency. |
Evidence clusters and their functions in the synthesis.
3.1.1 Direct clinical, instructional, and developmental supervision
The first cluster, direct clinical, instructional, and developmental supervision, provides the most explicit evidence base. examines supervision within principal evaluation in the UAE, showing that supervisory practice is closely tied to feedback, leadership evaluation, and accountability. examine instructional supervision in Oman during educational reform and show that supervisory roles are not static; they shift as reform expectations redefine what supervisors should do. focus specifically on supervisory feedback in principal evaluation, making feedback not a secondary issue but a central mechanism through which supervision can support formative improvement. examines teacher-supervisor relations in the Saudi TESOL context and highlights the importance of dialogue, classroom observation, post-observation conferencing, and power relations. add both supervisors’ and teachers’ perspectives from Saudi primary English education, making visible the gap that can emerge between support as intended by supervisors and support as experienced by teachers. extends the supervision literature into e-supervision readiness in Omani public schools, indicating that digital transformation introduces new technical, relational, and professional-capacity questions. directly examines developmental supervision approaches in Saudi Arabia, reinforcing the importance of differentiated support rather than uniform supervisory routines. contributes direct evidence on clinical supervision in Omani schools, while addresses developmental supervisory advancement and creative supervisory approaches in educational leadership.
Across these direct-supervision studies, three recurrent patterns emerge. First, supervision is consistently framed as an improvement-oriented practice, but its enactment is mediated by reform and evaluation pressures, hierarchical authority, and the quality of supervisory feedback (; , pp. 503–509; , pp. 237–238, 246; , pp. 1, 9). Second, supervision is not merely a technical sequence of observation and reporting; its developmental value depends on trust, role clarity, professional dialogue, differentiated support, and the supervisor's capacity to translate classroom evidence into actionable learning conversations (, pp. 64–65, 71; , pp. 503–509; , pp. 62–64; , pp. 1, 9). In technology-mediated supervision, this process also depends on adequate infrastructure, training, communication routines, and technical support (, pp. 63–69, 80). Third, developmental supervision is visible but unevenly realised. Although the studies emphasise professional growth, teachers and supervisors may still experience observation, feedback, or evaluation as compliance-oriented when institutional routines privilege documentation, ratings, workload demands, or accountability over reflective inquiry and sustained support (; , pp. 237–238, 246; , pp. 62–64; , pp. 1–2; ).
3.1.2 Teacher evaluation, appraisal, and data-driven assessment
The second cluster, teacher evaluation, appraisal, and data-driven teacher assessment, shows how supervision becomes formalized through standards, appraisal criteria, internal evaluators, teacher-led systems, and data tools. examine teachers’ experiences within the teacher evaluation process in UAE public schools, showing that teachers encounter evaluation through both formative and summative lenses. studies teacher evaluation in Saudi Arabia relative to national and international standards and best practices, foregrounding the need for alignment between evaluation criteria, feedback quality, evaluator roles, and reform expectations. examine teacher evaluation in Kuwait by heads of department, teachers themselves, peers, and students, expanding the evaluator role beyond the principal or external supervisor. reframes appraisal as professional development through teacher-led processes, including coaching, walk-throughs, reflective conversations, and teacher agency. introduce a data-driven teacher evaluation system in Oman using AHP-TOPSIS, indicating that evaluation is increasingly shaped by analytic and decision-support models. examines teacher evaluation policies and practices in Kuwaiti primary schools, while Teacher Evaluation in Kuwait (2016, pp. 1–3) evaluates the current Kuwaiti system and considers risk-based analysis as a principle for development. The Teacher Evaluation System in Saudi Arabia (2015, pp. 1–3) adds educators’ perceptions from female schools in one Saudi district.
The teacher-evaluation cluster demonstrates that appraisal is one of the most common institutional forms through which supervision is operationalized in GCC schools. Unlike clinical supervision, which ideally begins with teacher learning needs and observation-based dialogue, formal evaluation often begins with standards, rubrics, performance ratings, or institutional criteria. This does not make evaluation inherently anti-developmental. Several studies indicate that evaluation can support professional learning when teachers receive credible feedback, participate in reflective conversations, understand the criteria, and perceive evaluators as competent and fair (, pp. 1–2, 4, 14; ; ). However, the cluster also shows a persistent risk: evaluation can narrow supervision into a judgment event if appraisal emphasizes compliance, scores, or risk classification without sustained coaching. Multi-evaluator systems may broaden evidence and reduce dependence on one evaluator, but they also require calibration, trust, and clear evidence protocols (; , pp. 588–589). Data-driven evaluation tools can improve transparency and decision structure, but they require careful validation so that data systems do not replace professional interpretation (, pp. 192–194).
3.1.3 School inspection, institutional review, quality assurance, and accountability
The third cluster, school inspection, institutional review, quality assurance, and accountability, positions supervision at the system level. uses institutional review evidence to examine principals’ challenges in Bahraini government primary schools, especially in relation to leadership, school self-evaluation, professional development, and improvement. examine accountability and quality assurance for leadership and governance in Dubai's educational marketplace, demonstrating how external accountability shapes school leadership practice. study perceptions of the school self-evaluation process in Abu Dhabi, showing that internal review can become a bridge between external inspection and school-level improvement. links school leadership practices and policies with Abu Dhabi governmental school inspection outcomes. examines the impact of a quality assurance system on private education in Qatar from evaluators’ and principals’ perspectives. explicitly frames school inspection for school improvement through the leadership role. connects inspection reports with international assessment outcomes. examines the UAE school inspection framework as a transformation and performance-improvement tool. directly addresses inspection impact on teaching and learning.
This cluster shows that inspection and quality assurance function as macro-supervisory structures. They do not supervise teachers in the same way as a clinical supervisor observes a lesson and holds a post-observation conference, but they establish the performance language, evidence expectations, accountability pressure, and improvement priorities within which school-level supervision operates. The evidence suggests that inspection can stimulate improvement by making teaching quality, student outcomes, leadership, governance, and school self-evaluation visible (; , pp. 247–248; ). However, inspection also risks becoming detached from developmental supervision if schools respond mainly by producing documentation, preparing for visits, or managing ratings rather than strengthening everyday instructional practice. Studies of Dubai, Abu Dhabi, Qatar, Bahrain, and wider UAE inspection frameworks indicate that the developmental value of inspection depends heavily on leadership mediation: principals and middle leaders must translate external findings into internal coaching, targeted professional learning, and follow-up cycles (; ; ).
3.1.4 Instructional and middle leadership as developmental supervision
The fourth cluster, instructional leadership and middle leadership as developmental supervision, shows how supervision is enacted through leaders who may not always be formally titled supervisors. Vogel and Alhudithi (2023, pp. 1–3) examine Saudi and Qatari female principals’ preparation for and definitions of instructional leadership, including teacher supervision, classroom observation, feedback, and professional development. addresses principals’ difficulties in female Saudi secondary schools, including challenges linked to supervising, evaluating, and supporting teachers. examines distributed leadership applications in Kuwaiti high schools from teachers’ viewpoints, connecting supervision to teacher participation and school development. examine lesson observation in UAE schools as a tool for school improvement while contrasting accountability and perfunctory practice. studies the influence of heads of departments’ instructional leadership, cooperation, and administrative support on school-based professional learning in Kuwait. investigates instructional leadership in Kuwait's educational reform context from school leaders’ perspectives. links principal evaluation to instructional leadership behaviors. examines participatory lesson study as a way to strengthen and sustain English language teaching and leadership in Bahrain.
The instructional-leadership cluster expands the review's understanding of who supervises. In many schools, the practical work of supervision is carried out by principals, vice principals, heads of department, senior teachers, coordinators, or lesson-study facilitators. These actors influence teaching through lesson observation, feedback, curriculum monitoring, collaborative planning, peer learning, and professional development. The evidence indicates that leadership becomes developmental supervision when it is close to instruction, grounded in evidence of teaching and learning, and connected to sustained professional support rather than episodic checking (, ; Vogel and Alhudithi, 2023). At the same time, the literature also shows that instructional leadership can be constrained by workload, policy pressure, gendered organizational contexts, role ambiguity, and uneven preparation for feedback and classroom observation (; , pp. 124–127). Lesson observation is particularly revealing because it can operate either as meaningful inquiry into practice or as a perfunctory accountability event, depending on the quality of dialogue, follow-up, and trust (, pp. 222–225).
3.1.5 Coaching, mentoring, onboarding, and professional learning
The fifth cluster, coaching, mentoring, onboarding, and professional learning, represents the most explicitly developmental edge of the supervision continuum. examines instructional supervisors’ effectiveness in promoting personalized professional learning in private schools in Abu Dhabi, directly linking supervision with individualized professional growth. examines professional development of expatriate teachers in Kuwait through mentoring, showing how mentoring can support adaptation, retention, and professional identity. studies teacher onboarding as a multi-stage process framed by distributed leadership and ecological systems theory in American K-12 schools in the UAE, connecting induction, mentoring, professional support, and retention. directly examines coaching as a mechanism for changing teacher practice. identify mentoring and professional-development models relevant to supervision as capacity building. Badawy et al. (2025, p. 69) examines STREAM's impact on teaching, learning, and professionalism through teachers’ and lead teachers’ perspectives, making it relevant to teacher leadership, professional learning, and school improvement.
This cluster shows that supervision becomes most clearly developmental when it is connected to teacher learning needs, classroom practice, sustained support, and professional agency. Coaching and mentoring differ from inspection and formal evaluation because their authority is not primarily located in judgment but in support, modeling, questioning, and iterative improvement. However, the included studies also suggest that coaching and mentoring cannot succeed as isolated activities. They require time, leadership support, professional expertise, trust, and alignment with school improvement priorities (; ; , pp. 5–6, 348–349). In expatriate-heavy school systems, mentoring and onboarding also carry cultural and organizational significance because teachers may need support not only with pedagogy but with institutional expectations, curriculum language, parent communication, and accountability systems (, pp. 143–148; ). Professional development becomes supervisory when it is informed by evidence of practice and followed by feedback, coaching, or reflective inquiry rather than delivered as generic training (; Badawy et al., 2025).
3.2 Cross-cutting analytical themes
Across the 40 included studies, six analytical themes explain the current state of GCC educational supervision. These themes are supervision as an institutional continuum, feedback quality, role ambiguity and distributed supervision, inspection–development tension, data and technology, and professional-learning integration. Table 4 summarises the core interpretation, confidence level, and claims boundary associated with each theme.
Table 4
| Theme | Core interpretation | Confidence | Claims boundary |
|---|---|---|---|
| T1 Institutional continuum | Supervision ranges from compliance monitoring and inspection readiness to developmental capacity building. | Moderate | A synthesis pattern, not uniform implementation across every GCC system. |
| T2 Feedback quality | Developmental value depends on specific, timely, dialogic, evidence-based feedback and follow-up. | Moderate | Strongest in direct-feedback and evaluation studies; weaker in inspection-only evidence. |
| T3 Role ambiguity and distributed supervision | Supervisory functions are distributed across inspectors, leaders, middle leaders, mentors, coaches, peers, and teachers. | Moderate | Describes distributed functions; it does not imply identical formal roles across systems. |
| T4 Inspection–development tension | Inspection can stimulate improvement but can also generate performativity when disconnected from internal learning. | Moderate | Strongest in UAE/Dubai, Qatar, and Bahrain; GCC-wide claims require caution. |
| T5 Data and technology | E-supervision and data-driven evaluation are emerging and require ethical, contextual, and professional interpretation. | Moderate with boundaries | Based on emerging evidence, chiefly from Oman and Dubai-related studies. |
| T6 Professional-learning integration | Supervision becomes developmental when linked to coaching, mentoring, lesson study, onboarding, teacher-led appraisal, and sustained learning. | Moderate with boundaries | Several sources are indirect or context-specific; claims remain cautiously framed. |
Analytical themes, confidence, and claims calibration.
3.2.1 Theme 1: supervision as an institutional continuum
The first theme is supervision as an institutional continuum. At one end, supervision appears as monitoring, appraisal, standards checking, inspection, or institutional review, particularly in studies of teacher evaluation, principal evaluation, school inspection, self-evaluation, and quality assurance (; ; ; ; ; ; ). At the other, it appears as developmental supervision, instructional coaching, mentoring, lesson study, self- or peer-informed appraisal, personalized professional learning, and developmental feedback (; ; ; ; ; ; ; ). Most GCC studies are located between these poles, showing that supervision-related practice is rarely purely compliance-oriented or purely developmental (; ; ; ). The evidence therefore does not support a simple binary in which inspection is negative and coaching is positive. Instead, the key issue is whether supervisory mechanisms are connected to credible feedback, teacher agency, evidence use, and sustained professional learning (; ; ; ; ). Inspection can support improvement when it is translated into school self-evaluation, leadership action, reflective feedback, and internal improvement planning (; ; ; ). Appraisal can support growth when feedback is credible, specific, and dialogic (; ; ; ). Likewise, coaching and personalized professional learning remain developmental only when they include evidence, teacher voice, differentiation, and follow-up rather than reproducing top-down or episodic professional development (; ).
3.2.2 Theme 2: feedback quality
The second theme is the centrality of feedback quality. Feedback appears across direct supervision, principal evaluation, teacher evaluation, clinical and developmental supervision, instructional coaching, and lesson-observation studies (; ; ; ; ; ; ). The strongest evidence suggests that feedback is developmental when it is specific, evidence-based, timely, dialogic, and connected to next steps (; ; ; ). Feedback becomes weak when it is generic, delayed, primarily judgmental, episodic, or disconnected from teacher goals and professional-learning needs (; ; ; ). This finding aligns the GCC corpus with broader feedback research on actionable and dialogic feedback, but it also adds a regional insight: feedback is shaped by hierarchical expectations, evaluator authority, inspection pressure, workload, and cultural norms of professional communication (; ; ; ). Lesson-observation literature in the UAE further reinforces this point conceptually by warning that observation can become perfunctory when it is reduced to compliance evidence rather than used as a basis for professional dialogue and improvement ().
3.2.3 Theme 3: role ambiguity and distributed supervision
The third theme is role ambiguity and distributed supervision. Supervisors, principals, heads of department, inspectors, evaluators, mentors, and coaches may all influence teacher practice, but their responsibilities are not always clearly differentiated. Direct-supervision studies show that supervisory roles are shifting from inspectional or directive oversight toward more supportive, developmental, and reform-oriented practices, although older role expectations often persist (; ; ; ). Teacher-evaluation studies show that appraisal may involve principals, heads of department, peers, self-evaluation, students, and external or system-level evaluators, creating both opportunities and risks for multi-source evaluation (; ; ; ). Inspection and quality-assurance studies show that external review can influence internal leadership, self-evaluation, improvement planning, and school-level monitoring processes (; ; ; ; ). Instructional-leadership, coaching, and professional-learning studies further show that principals, heads of department, coaches, mentors, and senior or middle leaders may act as developmental supervisors even when their formal titles differ (; ; ; ; Vogel and Alhudithi, 2023). Role overlap can create productive distributed supervision, but only if actors share evidence standards, feedback protocols, and developmental purposes. Without such alignment, teachers may experience supervision as fragmented, inconsistent, or primarily compliance-oriented (; ; ).
3.2.4 Theme 4: inspection–development tension
The fourth theme is the inspection–development tension. Several GCC studies, particularly UAE-focused studies, show that external accountability can create urgency for improvement, clarify expectations, and provide structured evidence for school review, self-evaluation, and improvement planning (; ; ; ; ). However, inspection and observation systems can also generate performative or compliance-oriented behaviour when schools prioritise inspection readiness, documentation, or rating improvement over continuous instructional learning (; ; ). This tension is not unique to the GCC, but it is particularly visible in the corpus because inspection, quality assurance, school self-evaluation, and performance review have been prominent within several GCC reform systems, especially in UAE school-improvement policy (; ; ). The developmental challenge is therefore to ensure that inspection findings become inputs into internal improvement cycles, coaching, professional learning, middle-leadership development, and teacher inquiry rather than remaining isolated rating or compliance events (; ; ; ; ).
3.2.5 Theme 5: data and technology
The fifth theme is the emergence of data and technology as mediators of supervision, evaluation, and school improvement. In the GCC corpus, this appears in studies of e-supervision readiness in Oman, data-driven teacher evaluation systems, inspection-linked school performance analysis, school self-evaluation, and international assessment evidence (; ; ; ). shows that e-supervision depends not only on teachers’ and supervisors’ readiness but also on infrastructure, internet access, technical support, training, communication routines, and the perceived usefulness of digital supervision. extend this theme through TeacherEval, an AHP-TOPSIS system for teacher evaluation in Oman, showing how multi-criteria decision-making tools can structure appraisal criteria, weighting, and teacher ranking. demonstrates how Dubai private schools’ PISA, TIMSS, PBTS, and inspection-report data can be analysed together to examine school progress and inform policy and school-level decision-making. Data can therefore make supervision more transparent, systematic, and targeted, but the corpus also indicates that data does not interpret themselves. Data-driven supervision requires professional judgement, ethical use, contextual interpretation, and feedback literacy (; Badawy and Alkaabi, 2023; ; ). Without these safeguards, data systems risk narrowing teacher quality into measurable indicators that may not fully capture the relational, contextual, and pedagogical complexity of teaching (; Badawy and Alkaabi, 2023).
3.2.6 Theme 6: professional-learning integration
The sixth theme is professional-learning integration as the key condition that makes supervision developmental rather than merely procedural. The studies most strongly aligned with developmental supervision are those that connect observation, feedback, appraisal, coaching, mentoring, onboarding, lesson study, and instructional leadership to sustained teacher learning (; ; ; ; ; ). shows that instructional supervision becomes more meaningful when professional learning is personalised, job-embedded, and responsive to teacher voice and choice. similarly links instructional coaching, learning walks, instructional rounds, classroom observation, feedback, and follow-up to changes in teachers’ reported practice. and strengthen this theme through mentoring and professional-development evidence, while shows how lesson study can connect collaboration, reflective inquiry, and teacher learning. further shows that heads of department can support school-based professional learning when instructional leadership is combined with collegial cooperation and administrative support. This theme suggests that GCC supervision reform should not focus only on improving observation forms, appraisal rubrics, or inspection frameworks. It should build professional learning systems in which supervisory evidence leads to inquiry, collaboration, coaching, mentoring, targeted development, and follow-up. The central problem is therefore not the absence of supervision, but the uneven integration of supervision with sustained teacher professional learning.
3.3 GCC-sensitive conceptual model
The final synthesis can be expressed as a GCC-sensitive multi-level model of supervision. At the outer layer, national, emirate-level, and system-level reform agendas establish expectations through standards, inspection, accountability, quality assurance, principal evaluation, and school self-evaluation systems (; ; ; ; ). At the school level, principals and senior leaders translate these expectations into internal monitoring, teacher evaluation, lesson observation, school self-evaluation, improvement planning, and professional development (; ; ; ; ). At the middle-leadership level, heads of department, senior teachers, mentors, instructional supervisors, and coaches enact supervision through curriculum coordination, classroom observation, feedback conversations, lesson study, coaching, mentoring, and school-based professional learning (; ; ; ; ). At the teacher level, supervision is experienced as compliance, support, or professional agency depending on the quality of relationships, feedback, trust, teacher voice, and follow-up (; ; ; ). The model therefore positions developmental supervision not as a separate programme, but as the alignment of accountability, leadership, feedback, evidence use, and professional learning (; ; ; ). Figure 2 illustrates this multi-level alignment model.
Figure 2
3.4 Cross-contextual patterns and boundary conditions
Cross-contextual analysis identified nine related boundary conditions concerning country and sector variation, subject specificity, evaluator credibility, continuity, teacher agency, technology, evidence translation, student-learning linkages, and coherence among improvement mechanisms.
3.4.1 Country coverage and sector-specific supervisory conditions
A closer comparison of countries across the corpus shows that the supervision evidence base is unevenly distributed across mechanisms, sectors, and national contexts. The UAE appears prominently in studies of teacher evaluation, principal supervision, inspection, school self-evaluation, lesson observation, instructional coaching, personalized professional learning, and data-driven school improvement (; ; ; ; ; ; ; ; ; ). Oman appears strongly in direct instructional supervision, clinical supervision, e-supervision readiness, and data-driven teacher evaluation (; ; ; ). Saudi Arabia appears in studies of teacher-supervisor relations, developmental supervision, English-language teaching supervision, teacher evaluation, female school leadership, and instructional leadership (; ; ; ; ; ; Vogel and Alhudithi, 2023). Kuwait appears in teacher evaluation, distributed leadership, heads of department, school-based professional learning, and mentoring-based professional development (; ; ; ; ; ). Qatar appears in professional-development models and female principals’ understandings of instructional leadership (; Vogel and Alhudithi, 2023). Bahrain appears in institutional review, principal evaluation, and lesson-study-based professional learning (; ; ). This distribution indicates that GCC supervision research is not reducible to one mechanism, such as inspection or teacher evaluation alone. However, it also shows clear boundary conditions: some countries are represented through stronger direct-supervision evidence, while others are represented mainly through inspection, leadership, professional learning, or appraisal studies.
The corpus also suggests that public and private sectors create different supervisory conditions. Studies of public systems often emphasise reform mandates, ministry policies, inspection expectations, principal evaluation, teacher evaluation, and national accountability structures (; ; ; ; ; ). Studies of private or international schools more often highlight quality assurance, market-facing accountability, instructional coaching, personalised professional learning, mentoring, and school-level improvement processes (; ; ; ; ; ). This distinction does not mean that public systems are only compliance-oriented or that private systems are only developmental. Rather, it shows that sector context shapes the pressures, resources, autonomy, and accountability conditions under which supervision is enacted. Public systems may have stronger formal policy frameworks and ministry-directed evaluation structures, while private and international schools may have more varied internal models depending on ownership, curriculum, accreditation, inspection requirements, market pressures, and leadership capacity.
3.4.2 Subject specificity and supervisory expertise
Another pattern concerns subject specificity. Several studies are located in English-language teaching or TESOL contexts, including teacher–supervisor relations in Saudi TESOL, English teaching in Saudi primary schools, e-supervision readiness among English teachers in Oman, and participatory lesson study in Bahraini English-language teaching (; ; ; ). This concentration may reflect the importance of English in GCC reform agendas, the presence of expatriate teachers, and the visibility of English teaching as a site of observation, feedback, curriculum reform, and professional development. It also suggests that supervision should not be treated as entirely generic across subjects. Supervisory feedback in language teaching may require expertise in communicative pedagogy, curriculum standards, assessment, classroom interaction, and teacher language identity. By implication, subject-specific expertise is also likely to matter in mathematics, science, Arabic, Islamic education, early childhood education, and other fields, although the present corpus provides less direct subject-specific evidence for those areas. The evidence therefore supports the need for subject-sensitive supervision, especially through heads of department, senior teachers, coaches, mentors, and professional learning communities (; ; ; ).
3.4.3 Evaluator credibility and professional expertise
The corpus further reveals that teachers’ experiences of supervision are shaped by the perceived credibility of the supervisor or evaluator. Credibility is not only a matter of formal authority. It includes pedagogical expertise, fairness, consistency, ability to provide specific feedback, knowledge of curriculum, understanding of classroom context, and willingness to engage in dialogue. Studies of teacher evaluation, principal supervision, and teacher–supervisor relations show that feedback is more likely to be accepted when teachers believe the evaluator understands the work being evaluated and can translate evidence into meaningful professional guidance (; ; ; ). Studies of heads of department, instructional leadership, and personalised professional learning similarly indicate that middle leaders and supervisors are more effective when they combine subject or pedagogical knowledge, collaborative relationships, administrative support, and responsiveness to teachers’ professional needs (; ; ). This finding explains why formal evaluation tools alone cannot guarantee developmental supervision. The person interpreting and discussing the evidence matters.
3.4.4 Continuity, follow-up, and supervisory cycles
The studies also show that supervisory systems have temporal weaknesses. Observation and evaluation may occur at scheduled points, but teacher learning requires ongoing cycles. Inspection may occur periodically, but school improvement requires continuous internal review. Coaching may begin with an observed need, but practice change requires repeated feedback, implementation support, and follow-up. These temporal issues are visible in studies of teacher evaluation, instructional coaching, personalised professional learning, mentoring, school self-evaluation, and inspection follow-up (; ; ; ; ; ). The review therefore identifies continuity as a developmental condition. Supervision becomes weak when it is episodic, disconnected, or event-based; it becomes stronger when it forms a sustained cycle of evidence, feedback, professional learning, implementation, and review.
3.4.5 Teacher agency and enabling organisational conditions
The relationship between supervision and teacher agency is another important result. In compliance-oriented models, teachers are positioned mainly as objects of observation and evaluation. In developmental models, teachers participate in diagnosis, goal setting, reflection, experimentation, and evidence review. Several studies point toward more agentic forms of supervision, including peer and self-evaluation in teacher appraisal, lesson study, mentoring-based professional development, instructional coaching, and personalised professional learning (; ; ; ; ). These studies suggest that teacher agency is not opposed to accountability. Rather, agency can make accountability more meaningful when teachers are involved in interpreting evidence, identifying professional-learning needs, and designing improvement responses (; ; ).
At the same time, the evidence indicates that agency requires supportive structures. Teachers cannot participate meaningfully in appraisal, lesson study, coaching, or personalised professional learning if time, trust, facilitation, and leadership support are absent (; ; ). Mentoring-based professional development is also unlikely to support teachers effectively if mentoring is not structured, relational, and responsive to teachers’ professional and contextual needs (; ). Coaching cannot change practice if it is disconnected from classroom evidence, feedback, implementation support, and follow-up (; ). Similarly, distributed leadership cannot support developmental supervision if heads of department and middle leaders lack time, authority, collegial cooperation, or administrative support (; ). Thus, developmental supervision requires both professional culture and organisational design. The studies collectively show that agency emerges when teachers are invited into structured professional processes, not when responsibility for improvement is simply shifted onto them.
3.4.6 Data and technology in supervision
The technology-related studies introduce a further result: digitalisation creates opportunities for more flexible supervision but also new risks. E-supervision readiness in Oman suggests that teachers and supervisors may see potential in technology-supported observation, communication, documentation, and feedback, but readiness depends on infrastructure, skills, attitudes, training, technical support, and policy clarity (). Data-driven teacher evaluation in Oman illustrates the potential for structured decision support, while also raising questions about indicator selection, weighting, fairness, transparency, and professional interpretation (). Inspection-linked performance analysis in Dubai private schools shows how PISA, TIMSS, PBTS, and inspection-report data can frame school-improvement conversations, but such data still require contextual interpretation rather than mechanical application (). The result is not a rejection of digital supervision. It is a call for data-informed, not data-determined, supervision, in which digital tools support professional judgement rather than replace it (; Badawy and Alkaabi, 2023; ).
3.4.7 Evidence translation across system levels
The synthesis also indicates that supervision quality is shaped by the level at which evidence is interpreted. At the system level, inspection, quality assurance, school self-evaluation, and international assessment evidence tend to aggregate performance patterns and identify school-wide priorities (; ; ; ; ). At the school level, leaders interpret these patterns in relation to staffing, curriculum, assessment, professional development, and improvement planning (; ; ). At the department or middle-leadership level, heads of department, senior teachers, instructional supervisors, mentors, and coaches translate broad priorities into subject-specific practices, coaching routines, lesson study, and school-based professional learning (; ; ; ). At the classroom level, teachers need feedback that connects directly to lesson design, pedagogy, assessment, student engagement, learner needs, and implementation follow-up (; ; ; ; ). Problems occur when evidence remains at too high a level. A school may know from inspection that teaching requires improvement, but teachers may not know which instructional moves need to change. Studies of school self-evaluation, inspection follow-up, lesson observation, teacher evaluation, coaching, and professional learning collectively suggest that developmental supervision requires evidence to move downward from institutional diagnosis to classroom-level action and upward again through follow-up evidence that shows whether practice has changed (; ; ; ; ).
3.4.8 Links between supervision and student learning
The included studies also differ in how they position student learning. Some sources connect supervision-related practices to teaching and learning through lesson observation, inspection, instructional coaching, school improvement, STREAM-related teaching professionalism, or performance-analysis frameworks (; ; Badawy et al., 2025; ; ; ). However, these links are usually conceptual, perceptual, procedural, or based on school-level performance analysis rather than direct causal measurement of student outcomes. Other studies focus more on teachers’ or leaders’ perceptions, evaluation processes, feedback quality, professional learning, or institutional systems without directly measuring student learning effects (; ; ; ; ). This pattern is important. It suggests that the current GCC evidence base is stronger in describing supervisory processes, stakeholder experiences, and institutional mechanisms than in demonstrating causal effects on student learning. Future studies should therefore trace the full improvement chain from supervision to teacher learning, instructional change, classroom conditions, and student outcomes.
3.4.9 Coherence of supervisory and improvement mechanisms
A final result concerns the language of improvement. Across the corpus, terms such as supervision, evaluation, inspection, appraisal, feedback, coaching, mentoring, professional learning, school improvement, accountability, instructional leadership, and teaching professionalism are sometimes used as if they naturally support one another. The synthesis shows that this assumption is unsafe. These mechanisms can support one another only when their purposes are coordinated. For example, appraisal can identify professional-learning needs, but it can also produce defensive compliance when feedback is superficial, episodic, or primarily judgemental (; ; ). Inspection can clarify improvement priorities, but it can also produce performative readiness or documentation rituals if external review is not translated into internal learning cycles (; ; ). Coaching, STREAM-related instructional improvement, and personalised professional learning can support classroom change, but they can also become disconnected from teacher voice, classroom evidence, follow-up, or school priorities if they are not embedded in coherent supervisory design (Badawy et al., 2025; ; ). The most defensible interpretation of the 40 studies is therefore that GCC systems do not need more supervisory activity in general; they need more coherent, evidence-informed, learning-oriented supervisory design.
4 Discussion
The formal synthesis generated six cross-cutting themes: supervision as an institutional continuum; feedback quality; role ambiguity and distributed supervision; the inspection–development tension; data and technology; and professional-learning integration. The Discussion interprets the relationships among these findings through five higher-order lenses. These lenses are not a second thematic analysis; they provide an explanatory structure for locating the six themes within wider scholarship on supervision, accountability, leadership, context, and teacher learning.
The overarching inference is that the developmental value of supervision is not determined by the presence or absence of observation, appraisal, inspection, data systems, coaching, or professional learning alone. It depends on how these mechanisms are coupled. Across the corpus, external accountability becomes developmentally productive when leaders translate system-level evidence into credible classroom-level diagnosis, dialogic feedback, differentiated support, professional learning, and follow-up (; ; ; ; ; ). The same mechanisms become compliance-oriented when evidence is used principally for rating, documentation, surveillance, or episodic judgement rather than sustained improvement (; ; ; ). This interpretation is supported most consistently by the institutional-continuum, feedback-quality, distributed-supervision, and inspection–development themes, where direct supervision, teacher evaluation, inspection, and leadership studies converge around the importance of role clarity, credible feedback, and developmental follow-through (; ; ; ; ).
Evidence for data and technology is promising but more bounded, because the available studies focus mainly on e-supervision readiness, data-driven teacher-evaluation design, inspection-linked performance analysis, and conceptual data-use risks rather than longitudinal evidence of digital supervision improving practice (; ; Badawy and Alkaabi, 2023; ). Professional-learning integration has moderate support, but its interpretation also requires caution because several contributing studies are concentrated in particular countries, sectors, or developmental mechanisms such as mentoring, lesson study, coaching, heads of department, and personalised professional learning (; ; ; ; ; ). The Discussion therefore distinguishes a defensible regional pattern from uniform implementation across all GCC systems.
4.1 Accountability-mediated supervisory practice
The first interpretive lens concerns the institutional mediation of supervision by inspection, appraisal, quality assurance, and performance accountability. Classical developmental supervision begins from the teacher–supervisor relationship and differentiates support according to professional need (; Sullivan and Glanz, 2013). The GCC evidence does not displace that logic, but it shows that the dyad is nested within a wider policy architecture. Direct-supervision studies from Oman and Saudi Arabia portray supervisors as improvement-oriented actors whose work is nevertheless shaped by reform mandates, administrative expectations, workload, and hierarchical authority (; ; ; ). Teacher-evaluation and principal-evaluation studies similarly show that developmental intent is mediated by standards, ratings, evaluator credibility, feedback quality, and the perceived consequences of judgement (; ; ; ; ). The continuum identified in the Results is therefore institutional rather than merely interpersonal: the same observation can serve inquiry, appraisal, inspection preparation, documentation, or several purposes simultaneously.
This conclusion is consistent with international inspection research that distinguishes intended improvement effects from strategic or performative responses. Comparative studies show that inspection can clarify expectations, stimulate self-evaluation, and create pressure for improvement, but can also encourage documentation rituals, narrowing, and temporary compliance when schools concentrate on the inspection event rather than the instructional problem (; ). analysis of panoptic performativity provides a strong caution against assuming that greater visibility automatically produces deeper learning. The GCC corpus contains analogous tensions, particularly in UAE-focused studies of Dubai and Abu Dhabi inspection, self-evaluation, quality assurance, lesson observation, and inspection-linked school improvement (; ; ; ; ). Bahrain evidence also contributes to this lens through institutional review and principal evaluation, although the mechanism is less inspection-dominant than in the UAE evidence (; ). The inspection-development theme is therefore strongest in UAE-focused evidence and is supported more selectively in Bahrain. It should be interpreted as a cross-contextual mechanism whose expression depends on system design, evaluation purpose, school capacity, and leadership mediation rather than as a uniform regional effect.
The critical issue is therefore not whether accountability should be present, but whether accountability evidence is converted into an internal improvement process. Inspection, appraisal, principal evaluation, and school self-evaluation can provide diagnosis, comparative standards, performance documentation, and public assurance (; ; ; ; ). They do not, by themselves, specify the classroom practices that should change or build the professional knowledge required to change them. Where schools lack subject-specific leadership, coaching expertise, protected time, or professional trust, external evidence may remain at the level of institutional classification or compliance documentation (; ; ; ). Where those capacities exist, the same evidence can be translated into department-level inquiry, observation-feedback cycles, coaching, mentoring, and monitored professional learning (; ; ; ). This explains why the review rejects a simple opposition between accountability and development: accountability can identify and legitimate priorities, while professional learning provides the mechanism through which those priorities are addressed.
Data and technology intensify this mediation. E-supervision, algorithmic teacher-evaluation models, performance dashboards, and assessment data may increase the volume, speed, and apparent precision of supervisory evidence. Yet data-based decision-making research emphasises an iterative process of goal definition, collection of multiple evidence types, sense-making, action, and evaluation rather than a direct movement from data to decision (). The limited GCC evidence on e-supervision and data-driven appraisal supports the potential for greater targeting, flexibility, and transparency, but it also raises questions about infrastructure, technical support, indicator validity, weighting, fairness, contextual interpretation, privacy, and human oversight (; ). Conceptual work on datafication similarly cautions that data-driven decision-making requires data quality, ethical use, interpretation, stakeholder involvement, and professional judgement rather than mechanical reliance on indicators (Badawy and Alkaabi, 2023). Accordingly, the manuscript's argument is not that digital supervision is inherently developmental or technocratic. Its contribution depends on whether data are embedded in professional sense-making and whether teachers can interrogate the evidence used to evaluate and support them.
4.2 Developmental supervision as an unevenly realised aspiration
The second interpretive lens concerns the gap between the developmental language of supervision and its uneven enactment. Across the corpus, supervision is frequently justified in terms of improvement, professional growth, or teacher support. However, developmental purpose is not established by terminology. It is realised through the quality of evidence, dialogue, goal setting, differentiated assistance, and follow-up. This distinction is visible in the contrast between direct-supervision studies that emphasise observation and conferencing, evaluation studies that combine formative and summative functions, and coaching, mentoring, lesson-study, and professional-learning studies that extend support across the teacher career cycle (; ; ; ; ; ; ). The strongest recurring mechanism is feedback quality, but the evidence also shows that feedback is vulnerable to role conflict, limited time, weak evaluator preparation, hierarchical expectations, workload, and uncertainty about whether the encounter is developmental or consequential (; ; ; ; ).
International feedback research strengthens this interpretation (). Feedback to teachers is most likely to support learning when it is task- or goal-directed, specific, credible, and connected to an opportunity to act (Thurlings et al., 2013). Large-scale teacher-evaluation research similarly shows that teachers may perceive evaluators as trustworthy, fair, and accurate while still reporting that evaluation feedback is not sufficiently useful for improving instruction; administrator training alone also appears insufficient when evaluators lack the time and skill required for frequent, high-quality feedback (). found that post-observation feedback often omitted critical feedback characteristics, and that goal setting was the only critical feedback characteristic associated with subsequent teacher performance. These findings support the GCC synthesis while qualifying any assumption that protocols, rubrics, or evaluator certification alone will make feedback developmental.
The GCC contribution is to show that feedback quality is simultaneously technical, relational, and institutional. Technical quality concerns the accuracy, specificity, and instructional relevance of the evidence. Relational quality concerns credibility, fairness, dialogue, trust, and the teacher's opportunity to explain classroom context. Institutional quality concerns the consequences attached to the conversation and whether formative coaching can be distinguished from summative judgement. This interpretation is consistent with Tuytens and Devos (2011), who show that leadership characteristics shape teachers' perceptions of the utility of evaluation feedback and their responses to it. Studies of supervisory feedback, teacher–supervisor relations, teacher evaluation, and teacher appraisal indicate that weaknesses in any one of these dimensions can limit teacher uptake of feedback (; ; ; ; ). Conceptual lesson-observation literature in the UAE reinforces this interpretation by warning that observation becomes less developmental when it is reduced to accountability evidence or perfunctory procedure rather than professional dialogue (). Feedback literacy in this review therefore refers not merely to giving comments, but to the shared capacity of supervisors and teachers to interpret evidence, test explanations, agree feasible next steps, and revisit implementation.
Developmental supervision also varies across career stage and workforce context. Mentoring and onboarding evidence from Kuwait and the UAE suggests that novice, expatriate, and newly appointed teachers may require support with institutional norms, curriculum expectations, professional identity, cultural orientation, and accountability systems in addition to pedagogy (; ). Coaching and personalised professional learning address a related but distinct need: iterative refinement of practice based on classroom evidence, teacher voice, differentiated support, and follow-up (; ). These forms should not be treated as interchangeable. The corpus supports differentiated supervision, but the strength of this conclusion is moderate and bounded because the evidence is concentrated in selected countries, sectors, and supervision-adjacent mechanisms such as mentoring, onboarding, coaching, lesson study, and personalised professional learning. The appropriate conclusion is that developmental models are visible and plausible in the region, not that they are institutionally established across all GCC systems.
4.3 Leadership as the mediating mechanism between policy and practice
The third interpretive lens identifies leadership mediation as the mechanism connecting policy expectations with classroom-level support. The corpus broadens the identity of the supervisor beyond formally designated supervisors to include principals, vice principals, heads of department, senior teachers, instructional supervisors, mentors, coaches, peers, and teacher leaders. This distribution is analytically important because system-level evidence is often too aggregated to guide changes in lesson design, questioning, differentiation, assessment, or subject pedagogy. Principals and senior leaders mediate policy expectations through school priorities, teacher evaluation, improvement planning, and instructional leadership (; ; ; Vogel and Alhudithi, 2023). Middle leaders occupy a further translation point at which school priorities can be converted into subject-specific diagnosis, feedback, collaborative planning, coaching, lesson study, and school-based professional learning (; ; ; ; ).
Distributed leadership theory is useful here because it locates leadership in interactions among people, tools, routines, and situations rather than in a single office (Spillane, 2006). The review nevertheless adds an important qualification: distribution is not necessarily developmental. When roles, purposes, and evidence standards are unclear, multiple evaluators can generate inconsistency, duplication, or fragmented accountability (; ; ). Productive distribution requires role architecture: external reviewers establish system-level judgements; senior leaders set priorities and resource improvement; middle leaders provide subject-specific guidance; coaches support deliberate practice; mentors support induction; peers contribute collaborative inquiry; and teachers participate in diagnosis and goal setting. These functions may overlap, but the reason for the overlap and the consequences of each interaction must be explicit.
This mediating role is constrained by capacity and organisational conditions. Studies of principal difficulty, principal supervision, female instructional leadership, heads of department, instructional coaching, personalised professional learning, and lesson observation indicate that leaders may be expected to enact multiple reforms without equivalent preparation, time, authority, role clarity, or access to coaching expertise (; ; ; ; ; ; Vogel and Alhudithi, 2023). Conceptual lesson-observation literature in the UAE similarly cautions that observation can become perfunctory when leadership systems emphasise accountability procedures without sufficient developmental capacity (). contextual leadership argument is therefore directly relevant: generic leadership practices must be adapted to institutional, sociocultural, political, and school-improvement conditions. The GCC-sensitive model treats leadership not as a neutral conduit but as an interpretive layer. Leaders decide which evidence matters, how external findings are framed, whose expertise is authorised, whether teachers experience feedback as judgement or inquiry, and whether follow-up is protected amid competing operational demands.
4.4 Institutional and sociocultural conditions shaping supervisory relationships
The fourth interpretive lens addresses context without reducing explanation to a generic notion of culture. The review found that hierarchy, communication norms, gendered organisational settings, expatriate employment, evaluator authority, and expectations of professional deference may influence supervisory relationships. These influences interact with policy pressure, evaluator preparation, workload, school sector, subject expertise, leadership capacity, and data systems (; ; ; ; Vogel and Alhudithi, 2023). argues that leadership research should distinguish institutional, community, sociocultural, political, economic, and improvement contexts rather than relegate context to a background variable. Applying that principle to the GCC evidence prevents deterministic claims, such as attributing limited dialogue, compliance, or feedback avoidance principally to “GCC culture.”
Several alternative explanations should remain visible. Limited feedback dialogue may reflect deference to authority, but it may also reflect high-stakes appraisal, uncertainty about confidentiality, insufficient evaluator expertise, weak feedback protocols, or lack of time (; ; ; ; ). Differences in instructional leadership in gender-segregated settings may be culturally mediated, but they may also arise from preparation pathways, staffing structures, workload, school size, role expectations, or access to professional networks (; Vogel and Alhudithi, 2023). Expatriate teachers may require cultural and institutional induction, yet their experiences also vary by school ownership, curriculum, contract conditions, prior international experience, language, and mentoring arrangements (; ). The available studies illuminate these mechanisms, but they do not permit a single regional cultural explanation.
Country and sector distributions further delimit the argument. The UAE contributes the broadest evidence across inspection, teacher evaluation, principal supervision, school self-evaluation, lesson observation, instructional coaching, personalised professional learning, and data-informed school improvement (; ; ; ; ; ; ; ; ; ). Oman is prominent in direct supervision, clinical supervision, e-supervision readiness, and data-driven teacher evaluation (; ; ; ). Kuwait contributes evidence on teacher evaluation, distributed leadership, heads of department, school-based professional learning, and mentoring-based professional development (; ; ; ; ; ). Saudi Arabia contributes evidence on direct supervision, developmental supervision, teacher–supervisor relations, teacher evaluation, English-language teaching supervision, female school leadership, and instructional leadership (; ; ; ; ; ; Vogel and Alhudithi, 2023). Bahrain and Qatar are represented through narrower clusters, including institutional review, principal evaluation, lesson study, professional development, and female principals’ instructional leadership (; ; ; ; Vogel and Alhudithi, 2023). Public-system studies often foreground ministry policy, inspection, principal evaluation, teacher evaluation, and formal accountability structures, whereas private or international-school studies more often examine market-facing accountability, coaching, personalised professional learning, mentoring, onboarding, and retention (; ; ; ; ; ). These are tendencies within the corpus rather than essential differences between sectors. The review therefore uses country- and sector-qualified claims where the evidence is concentrated and reserves GCC-wide language for mechanisms supported across multiple settings.
4.5 Supervision as professional learning and school improvement
The fifth interpretive lens positions professional learning as the integration mechanism that converts supervisory activity into instructional improvement. The review's strongest developmental examples connect observation, evaluation, or feedback evidence with instructional coaching, mentoring, peer and self-evaluation, lesson study, collaborative inquiry, school-based professional learning, or personalised professional learning (; ; ; ; ; ; ). This supports the argument that supervision is incomplete when it ends with an observation record, appraisal score, inspection judgement, or improvement recommendation. The critical question is what teachers are enabled to learn, practise, test, and review after the judgement or feedback event.
The wider professional-learning literature supports this emphasis while cautioning against checklist approaches. and identify features such as sustained duration, content focus, active learning, collaboration, and coherence. Meta-analytic evidence also indicates that instructional coaching can positively influence instructional practice and student achievement (). , however, argues that design features alone do not explain how professional development changes teaching; programmes embody different theories of action about what teachers need to learn and how new ideas enter existing practice. similarly conceptualise teacher learning as an interaction among the teacher, the school, and the learning activity. More recent meta-analytic work proposes that effective professional development combines mechanisms that develop insight, motivate change, build teaching techniques, and embed those techniques in practice (Sims et al., 2025). These perspectives reinforce the review's claim that supervision should be designed as a learning system rather than a sequence of disconnected events.
This interpretation also clarifies the role of evidence review. Developmental supervision should generate evidence of process—observation records, coaching goals, mentoring plans, teacher reflection, professional-learning targets, and implementation notes—and evidence of impact, such as changes in lesson design, classroom interaction, assessment practice, student engagement, or, where the research design permits, student outcomes. iterative model of data use underscores the need to move from goals and evidence collection through sense-making and action to evaluation. In the GCC corpus, that full chain is rarely traced. Studies of inspection, teacher evaluation, supervisory feedback, instructional coaching, mentoring, lesson study, and personalised professional learning provide useful evidence about processes, perceptions, and implementation conditions, but they less often demonstrate sustained instructional effects or causal effects on student learning (; ; ; ; ; ; ; ). Performance-analysis studies can connect school improvement to broader student-assessment indicators, but they do not establish that supervision itself caused those outcomes (). Consequently, the review can defend a theory of how supervision may support improvement, but not a causal claim that any particular regional model improves student learning.
The resulting theoretical contribution is a multi-level alignment model. At the policy layer, standards, inspection, quality assurance, principal evaluation, and school self-evaluation define expectations and produce system-level evidence (; ; ; ; ). At the organisational layer, school leaders interpret priorities, allocate time and expertise, and establish the relationship between evaluation and development (; ; ; ; Vogel and Alhudithi, 2023). At the practice layer, middle leaders, instructional supervisors, coaches, mentors, and peers translate priorities into subject- and teacher-specific support (; ; ; ; ; ). At the centre, teachers interpret evidence, exercise agency, and test changes in practice through feedback, peer and self-evaluation, reflection, coaching, mentoring, and professional learning (; ; ; ). Feedback quality, role clarity, professional trust, evidence literacy, and data sense-making operate as cross-layer conditions. The model does not assume linear causality: weak alignment at any layer can interrupt the improvement pathway, while reciprocal evidence from classrooms should inform school and system review.
This expanded account is internationally relevant because it connects the classical supervisor–teacher dyad with the policy and organisational conditions that shape it. Clinical cycles of pre-observation, observation, analysis, conference, and follow-up remain valuable (; ), but the GCC evidence illustrates why those cycles cannot be analysed apart from inspection regimes, teacher-evaluation policies, leadership distribution, workforce diversity, data systems, and professional-learning infrastructure (; ; ; ; ; ; ; ). The review therefore extends supervision theory from a model of interpersonal assistance toward a context-responsive account of multi-level alignment. Its central proposition is not that accountability must be removed, but that accountability evidence requires leadership mediation, feedback literacy, evidence sense-making, and professional-learning capacity if it is to contribute to instructional improvement rather than end in classification or compliance (; ; ; ; ).
5 Recommendations
5.1 Recommendations for research
5.1.1 Trace the full improvement pathway
Future studies should test the sequence proposed by the conceptual model: supervisory evidence, teacher interpretation, professional learning response, change in instructional practice, and learner outcomes. Longitudinal, mixed-methods, and implementation designs should combine observations, feedback artefacts, coaching or mentoring records, teacher-learning evidence, and appropriately selected student indicators. This would move the field beyond perception-only accounts and identify where the proposed pathway is interrupted.
5.1.2 Compare systems, sectors, and underrepresented contexts
Comparative research should examine variation across GCC countries, public and private sectors, national and international curricula, school phases, gendered organisational settings, and urban or rural locations. Bahrain and Qatar require broader empirical coverage, while current evidence from the UAE, Oman, Saudi Arabia, and Kuwait should be tested beyond the clusters in which it is concentrated. Comparisons should investigate mechanisms and boundary conditions rather than assume that country labels explain differences.
5.1.3 Study feedback as a relational and institutional process
Research should examine the content, credibility, timing, dialogue, goal setting, and follow-up of supervisory feedback, together with the consequences attached to the interaction. Studies should distinguish formative coaching from summative evaluation and include both supervisor and teacher perspectives, direct observation of conferences, and analysis of written feedback. Particular attention is needed to how power, trust, role clarity, subject expertise, and evaluator workload shape feedback uptake.
5.1.4 Test middle-leadership and professional-learning mechanisms
Heads of faculty, heads of department, senior or lead teachers, coaches, and mentors should be examined as mediators between system priorities and classroom change. Quasi-experimental, longitudinal, and comparative case designs can test whether preparation in observation, coaching, data use, adult learning, and subject-specific pedagogy changes the quality of supervision and professional learning. Research should also compare teacher-led appraisal, lesson study, coaching, mentoring, and onboarding rather than treating them as a single developmental category.
5.1.5 Evaluate digital and data-mediated supervision
Emerging e-supervision and data-driven evaluation models require validation of indicators, weighting decisions, reliability, transparency, privacy, bias, and human oversight. Research should investigate whether digital systems improve diagnosis and follow-up or merely increase documentation and surveillance. Teacher and leader data literacy, opportunities for professional sense-making, and mechanisms for challenging or contextualising automated judgements should be treated as core implementation outcomes.
5.1.6 Validate and refine the GCC-sensitive model
The proposed multi-level alignment model should be treated as a set of testable propositions rather than a settled regional theory. Future studies should examine whether feedback quality, role clarity, middle-leadership capacity, and professional-learning integration mediate the relationship between accountability evidence and instructional change; whether these relationships vary by context; and whether reciprocal classroom evidence informs school and system-level review.
5.2 Recommendations for policy and practice
Policy and practice recommendations should be proportionate to the confidence of the supporting evidence. Table 5 links each policy and practice recommendation to the responsible actors, proposed implementation mechanism, intended outcome, and strength of the supporting evidence.
Table 5
| Evidence-based priority | Responsible actors | Recommended action | Implementation mechanism and intended result | Evidence strength |
|---|---|---|---|---|
| Connect accountability to learning | Ministries, regulators, school governing bodies, senior leaders | Require inspection and evaluation findings to generate specific professional-learning and follow-up cycles. | Define an evidence flow from external diagnosis to school priorities, department support, teacher action, and monitored review; reduce isolated rating events. | Moderate |
| Separate developmental and consequential functions | Ministries, HR/evaluation authorities, school leaders | Clarify when supervision is formative, summative, diagnostic, coaching-oriented, or inspection-related. | Use explicit purpose statements, differentiated records, confidentiality rules, and—where feasible—separate coaching from high-stakes judgement; strengthen trust and candour. | Moderate |
| Clarify distributed supervisory roles | Systems and schools | Specify the complementary functions of inspectors, principals, vice principals, heads of department, coaches, mentors, peers, and teachers. | Use responsibility matrices, shared evidence standards, referral pathways, and escalation protocols; reduce duplication and inconsistent feedback. | Moderate |
| Build feedback and coaching capability | Leadership-development providers, schools, professional-learning teams | Develop diagnostic, dialogic, goal-setting, subject-specific, and follow-up expertise. | Use calibrated observation practice, analysis of authentic evidence, coached feedback rehearsals, and review of implementation; improve feedback usefulness rather than compliance with forms. | Moderate |
| Strengthen middle-leadership capacity | Ministries, school groups, principals | Prepare heads of department and senior teachers as subject-specific developmental supervisors. | Protect time, provide authority and training, and link department inquiry to school priorities; improve translation from system evidence to classroom practice. | Moderate |
| Make professional learning the completion of supervision | School leaders, middle leaders, coaches, mentors | Ensure every significant supervisory diagnosis leads to differentiated learning and review. | Connect observation and appraisal to coaching, mentoring, lesson study, collaborative inquiry, or targeted development, with subsequent evidence of implementation. | Moderate with boundaries |
| Govern data and digital supervision ethically | Ministries, technology providers, regulators, schools | Establish validity, transparency, privacy, contestability, and human-oversight requirements. | Audit indicators and algorithms, document weighting, train users in sense-making, and permit contextual explanation; support inquiry rather than data-determined judgement. | Moderate with boundaries |
| Adapt a common framework to context | Policy makers, school groups, school leaders | Use common principles with flexible implementation pathways rather than a single rigid regional model. | Retain purpose clarity, evidence integrity, dialogue, differentiated support, follow-up, and system alignment while adapting roles and routines to country, sector, phase, subject, and workforce context. | Moderate |
Evidence-linked recommendations for policy and practice.
The recommended direction is learning-oriented accountability rather than the removal of standards, inspection, or evaluation. Its feasibility depends on implementation capacity. Policies that require coaching, evidence review, or professional learning without protected time, trained personnel, role clarity, and access to subject expertise risk adding another compliance layer. Implementation should therefore be staged, resourced, and evaluated through evidence of teacher learning and changed practice, not only completion rates or documentation.
6 Conclusion
6.1 Answers to the research questions
Research Question 1. “How is educational supervision conceptualized and enacted across the GCC studies included in the review?”
Educational supervision is conceptualised and enacted across the GCC corpus as a multi-level institutional continuum rather than a single role or technique. It includes direct clinical, instructional, developmental, and e-supervision; formal teacher and leader evaluation; system-level inspection and quality assurance; instructional and middle leadership; and developmental mechanisms such as coaching, mentoring, onboarding, lesson study, and teacher-led appraisal. The evidence is strongest for the interaction of direct supervision, evaluation, inspection, and leadership, while some professional-learning and technology practices remain emerging.
Research Question 2. “How do supervision, inspection, accountability, teacher appraisal, and professional learning interact in GCC education systems?”
Supervision, inspection, accountability, appraisal, and professional learning interact through processes of mediation and translation. External and formal mechanisms can clarify expectations and generate evidence, but they become developmentally consequential only when leaders convert that evidence into credible diagnosis, dialogic feedback, differentiated support, and follow-up. When the mechanisms are weakly connected, the same architecture can produce ratings, documentation, role conflict, or performative compliance without sustained instructional learning.
Research Question 3. “What evidence exists regarding clinical, developmental, instructional, and inspection-oriented forms of supervision in GCC contexts?”
The corpus contains direct evidence of instructional, clinical, developmental, supervisory-feedback, and e-supervision practices, together with semi-direct evidence from appraisal and contextual evidence from inspection and quality assurance. Fully developed developmental models are visible but unevenly institutionalised. Accordingly, the review supports the existence of a developmental aspiration and a set of promising mechanisms, not a claim that any single supervision model is consistently implemented or effective across GCC systems.
Research Question 4. “What institutional, cultural, policy, and leadership conditions enable or constrain developmental and improvement-oriented supervision?”
Developmental supervision is enabled by role clarity, evaluator credibility, subject and instructional expertise, protected professional dialogue, teacher agency, middle-leadership capacity, ethical evidence use, organisational trust, and connection to sustained professional learning. It is constrained by high-stakes ambiguity, inspection pressure, episodic observation, workload, weak preparation, fragmented roles, limited follow-up, and data systems that privilege measurement over interpretation. These conditions are context-sensitive and vary across countries, sectors, school phases, subjects, gendered settings, and expatriate-intensive workforces.
Research Question 5. “How can the GCC evidence base contribute to a context-sensitive conceptual model of supervision beyond inspection?”
The GCC evidence contributes a context-sensitive model of supervision as multi-level alignment. Policy accountability and quality assurance form the outer layer; school leadership and internal evaluation interpret system expectations; middle leadership, coaching, and mentoring translate priorities into practice; and teacher professional growth is the central intended outcome. Feedback quality, role clarity, trust, evidence sense-making, and professional-learning integration condition the relationships among these layers. The model is interpretive and reciprocal rather than a proven linear causal chain.
6.2 Principal contribution and a concluding note
The review's principal contribution is to move beyond both the narrow supervisor-teacher dyad. GCC school systems provide an internationally relevant case of how developmental aspirations are reframed within centralised, reform-intensive, accountability-oriented, and professionally diverse environments. The synthesis suggests that systems do not necessarily need more supervisory activity; they need greater coherence among the purposes, evidence, actors, conversations, learning responses, and follow-up that constitute supervision.
The conclusion remains calibrated to the evidence. The review identifies recurrent mechanisms, tensions, and boundary conditions, but the available literature is more persuasive about supervisory processes and experiences than causal effects on teaching or student learning. Future empirical work should therefore test the proposed alignment model and trace the full pathway from accountability evidence to teacher learning, instructional change, and learner outcomes. Until then, the strongest defensible policy principle is that accountability evidence should be used through, rather than instead of, professional judgement and sustained teacher learning.
Statements
Data availability statement
The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.
Author contributions
HeB: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing. TA: Data curation, Formal analysis, Investigation, Writing – original draft, Writing – review & editing. HoB: Data curation, Formal analysis, Investigation, Writing – original draft, Writing – review & editing. SA-Z: Data curation, Formal analysis, Investigation, Writing – original draft, Writing – review & editing. MA-H: Data curation, Formal analysis, Investigation, Writing – original draft, Writing – review & editing.
Funding
The author(s) declared that financial support was not received for this work and/or its publication.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/feduc.2026.1892041/full#supplementary-material
References
1
AbdallahR. K.Al MaktoumS. B.Al MansooriM. K. (2023). “The road to lesson observation as a tool to school improvement: accountability vs. Perfunctory,” in Restructuring Leadership for School Improvement and Reform, eds. AbdallahA. K.AlkaabiA. M. (Hershey, PA: IGI Global), 222–252. 10.4018/978-1-6684-7818-9.ch012
2
AbdalregalH. A. A. (2015) The teacher evaluation system in Saudi Arabia: a descriptive study of educators’ perceptions in the female schools in one Saudi Arabian school district. Doctoral dissertation. Saint Louis University. ProQuest Dissertations and Theses Global. UMI No. 3715704.
3
Abdul-RaheemT. M. Y. (2025) Exploring professional development of expatriate teachers in Kuwait through mentoring. Doctoral dissertation. University of South Carolina. Scholar Commons.
4
Abu-TinehA. M.SadiqH. M. (2018). Characteristics and models of effective professional development: the case of school teachers in Qatar. Prof. Dev. Educ.44 (2), 311–322. 10.1080/19415257.2017.1306788
5
AbualkishikA. M.MostafaN. N.AlbriekiA.AbualkeshekE. (2025). Teachereval: an AHP-TOPSIS system for data-driven teacher evaluation in the sultanate of Oman. Int. J. Adv. Appl. Sci.12 (12), 192–206. 10.21833/ijaas.2025.12.017
6
AchesonK. A.GallM. D. (2015). Clinical Supervision and Teacher Development: Preservice and in-service applications. 6th ed.Hoboken, NJ: John Wiley Sons.
7
Al MaktoumS. B.Al KaabiA. M. (2024). Exploring teachers’ experiences within the teacher evaluation process: a qualitative multi-case study. Cogent Educ.11 (1), 2287931. 10.1080/2331186X.2023.2287931
8
Al ThehliS. M. A. (2023) An investigation into the impact of school leadership practices and school policies on Abu Dhabi governmental school inspection outcomes. Doctoral thesis. The British University in Dubai.
9
Al-AbdaliF. M. A. (2015) Clinical supervision in omani schools: practices and implementation. Master’s thesis. Sultan Qaboos University.
10
Al-HadhramiM. A. S. (2023) Exploring English teachers’ and supervisors’ readiness towards the future implementation of e-supervision in al hamra and bahla public schools in al dakhiliya governorate. Master’s thesis. Sultan Qaboos University.
11
Al-KiyumiA.HammadW. (2019). Instructional supervision in the sultanate of Oman: shifting roles and practices in a stage of educational reform. Int. J. Leadersh. Educ.22 (2), 237–249. 10.1080/13603124.2018.1543570
12
AldaihaniS. G. (2020). Distributed leadership applications in high schools in the state of Kuwait from teachers’ viewpoints. Int. J. Leadersh. Educ.23 (3), 355–370. 10.1080/13603124.2018.1562096
13
AljenahiN. B. (2016) Teacher evaluation policies and practices in Kuwaiti primary schools. Doctoral thesis. Newcastle University.
14
AlkaabiA. M. (2025). A qualitative multi-case study of supervision in the principal evaluation process in the United Arab Emirates. Int. J. Leadersh. Educ.28 (1), 53–80. 10.1080/13603124.2021.2000032
15
AlkaabiA. M.AlmaamariS. A. (2020). Supervisory feedback in the principal evaluation process. Int. J. Eval. Res. Educ.9 (3), 503–509. 10.11591/ijere.v9i3.20504
16
AlkutichM. E. (2015) Examining the impact of school inspection on teaching and learning: dubai private schools as a case study. Master’s dissertation. The British University in Dubai.
17
AlmutairiT. S. S. S. A. (2016) Teacher evaluation in Kuwait: evaluation of the current system and consideration of risk-based analysis as a principle for further development. Doctoral thesis. Durham University. Durham E-Theses.
18
AlmutairiT. S.ShraidN. S. (2021). Teacher evaluation by different internal evaluators: head of departments, teachers themselves, peers and students. Int. J. Eval. Res. Educ.10 (2), 588–596. 10.11591/ijere.v10i2.20838
19
AlsalehA. (2019). Investigating instructional leadership in Kuwait’s educational reform context: school leaders’ perspectives. Sch. Lead. Manag.39 (1), 96–12010.1080/13632434.2018.1467888
20
AlsalehA. A. (2022). The influence of heads of departments’ instructional leadership, cooperation, and administrative support on school-based professional learning in Kuwait. Educ. Manag. Adm. Leadersh.50 (5), 832–850. 10.1177/1741143220953597
21
AlshehriR. A. (2018) The practice of developmental supervision approaches in Saudi Arabia. Doctoral dissertation. Southern Illinois University Carbondale. ProQuest Dissertations & Theses Global. ProQuest No. 10811835.
22
AlsuhaibaniY.AltalhabS.BorgS.AlharbiR. (2023). 15 Years’ experience of teaching English in Saudi primary schools: supervisors’ and teachers’ perspectives. Linguist. Educ.77, 101222. 10.1016/j.linged.2023.101222
23
AlwadiH. M.MohamedN.WilsonA. (2020). From experienced to professional practitioners: a participatory lesson study approach to strengthen and sustain English language teaching and leadership. Int. J. Lesson Learn. Stud.9 (4), 333–349. 10.1108/IJLLS-10-2019-0072
24
AlzamilJ. (2021). Principals’ difficulties at female Saudi secondary schools. J. Educ. Learn.10 (2), 124–128. 10.5539/jel.v10n2p124
25
BadawyH. R. I.AlkaabiA. M. (2023). “From datafication to school improvement: the promise and perils of data-driven decision making,”. in Restructuring Leadership for School Improvement and Reform, eds A. K. Abdallah, and A. M. Alkaabi (Hershey, PA: IGI Global), 301–325, 10.4018/978-1-6684-7818-9.ch015
26
BadawyH. R. I.AlkaabiA. M.MohsenW. A.AlblooshiK. M. (2024). Developmental supervisory advancements: refining the art of crafting creative approaches and applications in educational leadership. Cutting-Edge Innovations in Teaching, Leadership, Technology, and Assessment, eds A. K. Abdallah, A. M. Alkaabi, and R. Al-Riyami (Hershey, PA: IGI Global), 100–119, 10.4018/979-8-3693-0880-6.ch008
27
BadawyH. R. I.AlkaabiA. M.QablanA.AbdallahA.TairabH.BadawyH. R. I. (2025). The impact of STREAM on teaching, learning, and professionalism: investigating teachers' and lead teachers' perspectives. SN Soc. Sci.5, 69. 10.1007/s43545-025-01104-x
28
BahlouliA. H. (2018) The impact of instructional coaching on changing teachers’ practice in three private schools in sharjah - UAE. Master’s dissertation. The British University in Dubai.
29
BarbourD. T. (2019) Bridging the gap: school inspection for school improvement, the leadership role. Master’s dissertation. The British University in Dubai.
30
BarnawiO. Z. (2016). Dialogic investigations of teacher-supervisor relations in the TESOL landscape. Cogent Educ.3 (1), 1217818. 10.1080/2331186X.2016.1217818
31
BdeirR. (2019) Investigating the progress of dubai private schools’ PISA and TIMSS results and school inspection reports from 2011 to 2018. Doctoral thesis. The British University in Dubai.
32
Ben JaafarS.AlzouebiK.BodolicaV. (2022). Accountability and quality assurance for leadership and governance in dubai-based educational marketplace. Int. J. Educ. Manag.36 (5), 641–660. 10.1108/IJEM-11-2021-0439
33
Blaik HouraniR.LitzD. (2016). Perceptions of the school self-evaluation process: the case of Abu Dhabi. Sch. Lead. Manag.36 (3), 247–270. 10.1080/13632434.2016.1247046
34
BrewerD. J.AugustineC. H.ZellmanG. L.RyanG.GoldmanC. A.StaszC.et al (2007). Education for a new era: Design and Implementation of K-12 Education Reform in Qatar. Santa Monica, CA: RAND Corporation.
35
ButiM. (2019) The perceptions of elementary school principals in Bahrain on the principal evaluation system and its impact on instructional behaviors. Doctoral dissertation. George Mason University. ProQuest Dissertations & Theses Global. ProQuest No. 27663682.
36
CoganM. L. (1973). Clinical Supervision. Houghton Mifflin.
37
Darling-HammondL. (2013). Getting Teacher Evaluation Right: What Really Matters for Effectiveness and Improvement. Teachers College Press.
38
Darling-HammondL.HylerM. E.GardnerM. (2017). Effective Teacher Professional Development. Learning Policy Institute.
39
DesimoneL. M. (2009). Improving impact studies of teachers’ professional development: toward better conceptualizations and measures. Educ. Res.38 (3), 181–199. 10.3102/0013189X08331140
40
EhrenM. C.GustafssonJ. E.AltrichterH.SkedsmoG.KemethoferD.HuberS. G. (2015). Comparing effects and side effects of different school inspection systems across Europe. Comp. Educ.51 (3), 375–400. 10.1080/03050068.2015.1045769
41
El MamloukS. G. (2023) The effectiveness of instructional supervisors in promoting personalized professional learning at four private schools in Abu Dhabi. Doctoral thesis. The British University in Dubai.
42
El SaadiD. H. (2017) The contribution of the UAE school inspection framework as a quality assurance tool for school transformation and performance improvement. Master’s dissertation. The British University in Dubai.
43
FateelM. J. (2025). Principals’ challenges in Bahraini government primary schools as reflected in institutional reviews: leadership aspect. Int. Educ. Stud.18 (6), 1–7. 10.5539/ies.v18n6p1
44
GlickmanC. D.GordonS. P.Ross-GordonJ. M. (2018). SuperVision and Instructional Leadership: A Developmental Approach. 10th ed.Pearson.
45
GoldhammerR. (1969). Clinical Supervision: Special Methods for the Supervision of Teachers. Holt, Rinehart and Winston.
46
GonzalezG.KarolyL. A.ConstantL.SalemH.GoldmanC. A. (2008). Facing Human Capital Challenges of the 21st Century: Education and Labor Market Initiatives in Lebanon, Oman, Qatar, and the United Arab Emirates. RAND Corporation.
47
GordonS. P. (2019). Educational supervision: reflections on its past, present, and future. Journal of Educational Supervision2 (2), 27–52. 10.31045/jes.2.2.3
48
Gulf Cooperation Council (1981). Charter of the Cooperation Council for the Arab States of the Gulf. Riyadh: General Secretariat of the Gulf Cooperation Council. Available online at:https://www.gcc-sg.org/en/AboutUs/Pages/PrimaryLaw.aspx (Accessed April 8, 2026).
49
GustafssonJ.-E.EhrenM. C. M.ConynghamG.McNamaraG.AltrichterH.O’HaraJ. (2015). From inspection to quality: ways in which school inspection influences change in schools. Stud. Educ. Eval.47, 47–57. 10.1016/j.stueduc.2015.07.002
50
GuthrieA. A. (2025) Supporting transition and enhancing retention: teacher onboarding as a multi-stage process framed by distributed leadership and ecological systems theory in American K-12 schools in the United Arab Emirates. Doctoral dissertation. George Mason University. ProQuest Dissertations & Theses Global.
51
HallingerP. (2005). Instructional leadership and the school principal: a passing fancy that refuses to fade away. Leadersh. Policy. Sch.4 (3), 221–239. 10.1080/15700760500244793
52
HallingerP. (2018). Bringing context out of the shadows of leadership. Educ. Manag. Adm. Leadersh.46 (1), 5–24. 10.1177/1741143216670652
53
HattieJ.TimperleyH. (2007). The power of feedback. Rev. Educ. Res.77 (1), 81–112. 10.3102/003465430298487
54
HewittM. (2024). “Professional development through teacher-led appraisal,” in Igniting Excellence in Faculty Development at International Schools eds. P. Pelonis, and T. Zaharopoulos, Cham: Palgrave Macmillan, 165–173. 10.1007/978-3-031-67055-8_9
55
HongQ. N.FàbreguesS.BartlettG.BoardmanF.CargoM.DagenaisP.et al (2018). The mixed methods appraisal tool (MMAT) version 2018 for information professionals and researchers. Educ. Inf.34, 285–291. 10.3233/EFI-180221
56
HunterS. B.SpringerM. G. (2022). Critical feedback characteristics, teacher human capital, and early-career teacher performance: a mixed-methods analysis. Educ. Eval. Policy. Anal.44 (3), 380–403. 10.3102/01623737211062913
57
KattanS. A. (2022) Teacher evaluation in Saudi Arabia relative to national and international teacher evaluation standards and best practices. Doctoral dissertation. Western Michigan University. ScholarWorks at WMU.
58
KennedyM. M. (2016). How does professional development improve teaching?Rev. Educ. Res.86 (4), 945–980. 10.3102/0034654315626800
59
KraftM. A.BlazarD.HoganD. (2018). The effect of teacher coaching on instruction and achievement: a meta-analysis of the causal evidence. Rev. Educ. Res.88 (4), 547–588. 10.3102/0034654318759268
60
KraftM. A.ChristianA. (2022). Can teacher evaluation systems produce high-quality feedback? An administrator training field experiment. Am. Educ. Res. J.59 (3), 500–537. 10.3102/00028312211024603
61
MohamedR. Y. (2021) The impact of a quality assurance system on private education in the state of Qatar: perspectives of evaluators and principals. Phd thesis. University of Warwick. Warwick Research Archive Portal.
62
OECD (2013). Teachers for the 21st Century: Using Evaluation to Improve Teaching. Paris: OECD Publishing. 10.1787/9789264193864-en
63
OECD (2024). Education Policy Outlook 2024: Reshaping Teaching into a Thriving Profession from ABCs to AI. Paris: OECD Publishing. 10.1787/dd5140e4-en
64
OpferV. D.PedderD. (2011). Conceptualizing teacher professional learning. Rev. Educ. Res.81 (3), 376–407. 10.3102/0034654311413609
65
PageM. J.McKenzieJ. E.BossuytP. M.BoutronI.HoffmannT. C.MulrowC. D.et al (2021). The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. Br. Med. J.372,, n71. 10.1136/bmj.n71
66
PerrymanJ. (2006). Panoptic performativity and school inspection regimes: disciplinary mechanisms and life under special measures. J. Educ. Policy21 (2), 147–161. 10.1080/02680930500500138
67
RobinsonV. M. J.LloydC. A.RoweK. J. (2008). The impact of leadership on student outcomes: an analysis of the differential effects of leadership types. Educ. Adm. Q.44 (5), 635–674. 10.1177/0013161X08321509
68
RomanowskiM. H.CherifM.Al AmmariB.Al AttiyahA. (2013). Qatar’s educational reform: the experiences and perceptions of principals, teachers and parents. Int. J. Educ. Dev.5 (3), 108–135. 10.5296/ije.v5i3.3995
69
SchildkampK. (2019). Data-based decision-making for school improvement: research insights and gaps. Educ. Res.61 (3), 257–273. 10.1080/00131881.2019.1625716
70
SimsS.Fletcher-WoodH.O’Mara-EvesA.CottinghamS.StansfieldC.GoodrichJ.et al (2025). Effective teacher professional development: new theory and a meta-analytic test. Rev. Educ. Res.95 (2), 213–254. 10.3102/00346543231217480
71
SpillaneJ. P. (2006). Distributed Leadership. San Francisco, CA: Jossey-Bass.
72
SullivanS.GlanzJ. (2013). Supervision That Improves Teaching and Learning: Strategies and Techniques. 4th ed.Thousand Oaks, CA: Corwin.
73
ThomasJ.HardenA. (2008). Methods for the thematic synthesis of qualitative research in systematic reviews. BMC. Med. Res. Methodol.8, 45. 10.1186/1471-2288-8-45
74
ThurlingsM.VermeulenM.BastiaensT.StijnenS. (2013). Understanding feedback: a learning theory perspective. Educ. Res. Rev.9, 1–15. 10.1016/j.edurev.2012.11.004
75
TimperleyH. (2008). Teacher Professional Learning and Development. Geneva: UNESCO International Bureau of Education.
76
TongA.FlemmingK.McInnesE.OliverS.CraigJ. (2012). Enhancing transparency in reporting the synthesis of qualitative research: eNTREQ. BMC. Med. Res. Methodol.12, 181. 10.1186/1471-2288-12-181
77
TuytensM.DevosG. (2011). Stimulating professional learning through teacher evaluation: an impossible task for the school leader?Teach. Teach. Educ.27 (5), 891–899. 10.1016/j.tate.2011.02.004
78
UNESCO (2019). Teacher Policy Development Guide. Paris: UNESCO.
79
UNESCO (2024). Global Education Monitoring Report 2024/5: Leadership in Education. Paris: UNESCO.
80
VogelL. R.AlhudithiA. (2023). Arab women as instructional leaders of schools: saudi and qatari female principals’ preparation for and definition of instructional leadership. Int. J. Leadersh. Educ.26 (5), 795–818. 10.1080/13603124.2020.1869310
Summary
Keywords
accountability, educational supervision, gulf cooperation council, inspection, instructional leadership, policy enactment, systematic review, teacher professional learning
Citation
Badawy HRI, Alshloul T, Badawy HRI, Al Zaabi SM and Al-Hasan M (2026) Educational supervision in GCC school systems is shaped by inspection, accountability, and context-sensitive instructional leadership: a systematic review. Front. Educ. 11:1892041. doi: 10.3389/feduc.2026.1892041
Received
26 May 2026
Revised
26 July 2026
Accepted
31 July 2026
Published
14 August 2026
Volume
11 - 2026
Edited by
Sana Al Maktoum, Zayed University, United Arab Emirates
Reviewed by
Annette Kappert, SRH University of Applied Sciences Heidelberg, Germany
William L. Sterrett, Baylor University, United States
Updates
Copyright
© 2026 Badawy, Alshloul, Badawy, Al Zaabi and Al-Hasan.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Hesham R. I. Badawy 202070281@uaeu.ac.ae
†These authors contributed equally to this work
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.