Abstract
Educational supervision has evolved internationally from a narrow inspection-and-compliance function toward a developmental practice supporting instructional leadership, professional learning, and sustained school improvement. This Policy Review examines UAE supervision through loose-coupling and instructional leadership theories. Using a structured narrative-synthesis methodology (Section 3), the review draws on 54 sources — 15 UAE-specific empirical or policy sources and 39 international theoretical, comparative, meta-analytic, or methodological sources — benchmarked against four contrasting international supervision models (Section 5). The review shows that the UAE's multi-authority governance structure (Ministry of Education, Abu Dhabi Department of Education and Knowledge, Knowledge and Human Development Authority) produces a loosely coupled accountability system that becomes progressively tighter, and more consequential for teaching quality, as it approaches the classroom. Current supervisory practice is dominated by observation-and-documentation cycles whose developmental value depends heavily on evaluator capacity, feedback quality, and psychological safety — moderators that international meta-analytic evidence identifies as the mechanisms through which supervision raises or fails to raise instructional quality. A critical synthesis finds UAE research concentrated among a relatively small number of overlapping teams and dominated by small-sample qualitative case studies, with regulator-published policy documents supplementing the peer-reviewed literature in ways that call for cautious generalization. Benchmarked internationally, the UAE emerges as an evolving hybrid model — sharing features with high-stakes European/English inspection, a design associated elsewhere with workforce-retention concerns, while also drawing on elements of Finland's trust-based supervision, and without yet incorporating a Singapore-style competency-track infrastructure for converting accountability data into developmental pathways. Current supervisory practice is dominated by observation-and-documentation cycles whose developmental value depends heavily on evaluator capacity, feedback quality, and psychological safety — moderators that international meta-analytic evidence identifies as the mechanisms through which supervision raises or fails to raise instructional quality. Building on this critique, the authors propose a conceptual model integrating governance, institutional, and classroom-level supervision, and identify four evidence-based innovation trajectories: coaching-oriented supervision (0.49 SD on instruction, 0.18 SD on achievement), AI-assisted instructional feedback, digitally mediated evidence systems, and cross-regulator coherence reform. Two comparative tables and a meta-analytic benchmark table render UAE and international evidence side by side. The contribution is fourfold: this is the first UAE supervision review to apply an explicit coupling-theory lens, make its methodology transparent, position the UAE against international models through international contextual benchmarking, and render its three-authority system as a single conceptual diagram for policymakers and school leaders.
1 Introduction
1.1 The contemporary evolution of educational supervision and leadership
Across international education systems, educational supervision has shifted decisively over the past three decades from compliance-led inspection toward a more developmental and leadership-centered function that strengthens instructional quality, professional learning, and sustained school improvement. Contemporary scholarship reinforces that supervision is most effective when it is embedded in instructional leadership routines — observation, evidence use, feedback, coaching, and professional dialogue — rather than treated as episodic monitoring (). The theoretical shift is now paralleled by robust empirical evidence: a meta-meta-analysis synthesizing twelve prior meta-analyses reports a statistically significant positive association between principal leadership and student achievement (Cohen's d = 0.34), with more recent and methodologically stronger studies producing more precise estimates (). A subsequent multivariate meta-analysis of 42 studies published between 2000 and 2020 confirms a similar magnitude (0.22–0.25 SD, depending on whether direct or indirect effects are modeled) and demonstrates that the observed effect is contingent on how principal impact is conceptualized (). A systematic review and meta-analysis of principal behaviors likewise reports direct effects on student achievement (0.08–0.16 SD), on teacher instructional practice (0.35 SD), on teacher well-being (0.34–0.38 SD), and on school organizational health (0.72–0.81 SD) (). Together, these syntheses justify treating supervision not as an administrative appendage but as a strategic instructional-leadership function whose design choices materially shape teaching and learning.
The syntheses cited above span several distinct evidentiary traditions that this review treats separately rather than as interchangeable evidence for “supervision.” Effects attributed to general school leadership (a principal's overall strategic and organizational influence) are not the same evidentiary category as effects attributed to instructional leadership (leadership actions specifically directed at classroom teaching), supervision (structured observation-feedback cycles), teacher evaluation (summative rating systems), external inspection (regulator-administered school review), or coaching (sustained, dialogic professional support). d = 0.34 estimate, for example, is a general-leadership effect and should not be read as a supervision-specific effect; Table 1 makes these distinctions explicit for every synthesis cited in this section, so that claims below are traceable to the correct evidentiary tradition.
Table 1
| Evidentiary category | What it measures | Key sources cited in §1.1 | Effect reported |
|---|---|---|---|
| General school leadership | Overall strategic/organizational influence of the principal on school outcomes | ; | d = 0.34; 0.22–0.25 SD |
| Instructional leadership | Leadership actions directed specifically at classroom teaching and learning | ; | 0.08–0.16 SD (student); 0.35 SD (practice) |
| Supervision (observation-feedback) | Structured, cyclical observation and feedback on classroom practice | 0.49 SD (instruction); 0.18 SD (achievement) | |
| Teacher evaluation | Summative rating systems, often value-added or observation-based | Identifies variation; no consistent effect on practice absent development | |
| External inspection | Regulator-administered school review with public reporting | ; | Mixed; workforce-retention costs documented |
| Coaching | Sustained, dialogic, practice-proximal professional support | ; | d = 0.41; moderated strongly by cognitive modeling (d = 0.90) |
Evidentiary categories distinguished in this review, mapped to the sources and effect sizes cited in §1.1 effects reported for one category (e.g., general leadership) should not be attributed to another (e.g., supervision).
Contemporary policy debates emphasize the persistent tension between supervision as summative accountability (ratings, compliance, consequences) and supervision as formative professional support (growth-focused feedback, coaching, and improvement planning). This tension is not a rhetorical framing but a wicked problem in the technical sense: elementary-principal case research shows that high-functioning leaders explicitly acknowledge — rather than resolve — the tensions and conflicts between supervision and evaluation, and that the productive move is to act as an instructional coach rather than a manager of teachers while holding both functions in view (). Reviews of teacher-evaluation reform in high-stakes accountability regimes converge on the same finding: value-added and observation-based ratings can identify variation in teacher effectiveness but, absent complementary developmental infrastructure, tend to produce compliance behavior rather than instructional change ().
1.2 Rationale, theoretical positioning, and contribution
Despite a growing body of UAE-based scholarship on evaluation and leadership, this literature remains theoretically under-specified: individual studies typically describe a single mechanism (e.g., principal portfolios, teacher-evaluation feedback) without situating it within a governance-level theory of why UAE supervision behaves the way it does. This review addresses that gap by adopting an explicit dual theoretical frame — loose-coupling theory and instructional leadership theory — introduced in Section 2 and applied consistently throughout the analysis.
The review makes four original contributions. First, it is, to our knowledge, the first synthesis to theorize the UAE's three-authority supervisory system (MOE, ADEK, KHDA) explicitly as a loosely-to-tightly coupled system, rather than describing the authorities in parallel. Second, it applies and reports a transparent, replicable review methodology (Section 3), moving the paper from a general narrative summary toward a structured Policy and Practice Review. Third, it translates the synthesis into a single conceptual diagram (Figure 1) and an explicit critical appraisal of the UAE evidence base (Sections 4.5, 5.6, 6.5, 7.6). Fourth, it triangulates UAE findings against international meta-analytic evidence on coaching, principal-leadership effects, and teacher-evaluation reform, allowing UAE-specific claims to be evaluated against the strongest available international benchmarks rather than accepted on their own terms. Collectively, these additions reposition the paper from a descriptive account of what UAE regulators and researchers have published toward an analytical account of how and why the system produces the outcomes documented.
Figure 1
2 Theoretical framework
This review is theoretically anchored in two complementary traditions, extended by the coupling-tiering framework developed in more recent institutional-theory scholarship.
2.1 Loose-coupling theory
foundational argument was that organizational elements in educational bureaucracies are tied together frequently and loosely rather than through the dense, tight linkages implied by classical bureaucratic theory. Weick argued that loose coupling incorporates a wide range of empirical observations about schools, and that the construct simultaneously creates methodological difficulties (couplings are hard to measure) and generates novel research questions. subsequently formalized loose coupling as a dialectical construct — organizations can be both distinctive and responsive — rather than a synonym for organizational slack.
Later institutional-theory work has extended this frame in two directions relevant to the UAE case. argued that the U.S. education system is moving toward fragmented centralization, in which environmental pressures, powerful new institutional actors, and institutional isomorphism together push policymakers to tighten coupling selectively rather than uniformly. , using seven waves of the U.S. Schools and Staffing Survey, demonstrated empirically that federal policy, district-school relationships, and individual principals each shape school-level coupling in distinct and partially independent ways — that is, coupling is a tiered phenomenon, not a single system-level property.
We use loose coupling — in this tiered form — as the framework for explaining a recurring pattern in the UAE evidence: national licensing standards, emirate-level inspection frameworks, and school-level practice are formally linked through shared vocabulary (e.g., “quality assurance,” “instructional leadership”) but are inconsistently coupled in implementation, producing the regulatory fragmentation, evaluator variability, and compliance-vs.-development tension documented across the reviewed studies.
2.2 Instructional leadership theory
Complementing this systems-level lens, instructional leadership theory (; ; ) provides the practice-level account of how supervision, when tightly coupled to classroom teaching through observation, feedback, and coaching, is associated with improved teacher efficacy and instructional quality. The strongest recent evidence for this claim comes from three independent lines of meta-analytic research. First, meta-analysis of 60 causal studies of teacher coaching found pooled effects of 0.49 SD on instruction and 0.18 SD on student achievement, though average effects in scaled-up effectiveness trials were substantially smaller than in tightly controlled efficacy trials. Second , meta-analysis of pre-service coaching, mentoring, and supervision studies reports a smaller but significant overall effect (d = 0.41), with cognitive modeling by supervisors emerging as a substantial moderator (d = 0.90). Third, process-level research on instructional coaching indicates that coach-teacher relationship quality predicts teacher engagement and reflection, while the quality of coaching strategies predicts overall classroom instructional quality (; ) — mechanistically specifying which coaching behaviors do the work.
A synthesis focused specifically on academic press, examining 79 quantitative studies over 30 years, reports a large effect of academic press on student achievement and a close-to-large effect of school leadership on academic press () — that is, leadership acts on student learning largely by shaping the school-level press for challenging instruction rather than by intervening directly in classrooms. A parallel meta-analysis of principal transformational leadership and teacher job satisfaction across ten studies reports a moderate correlation (r ≈ 0.50; ).
Read against Section 2.1, this practice-level evidence specifies what tight coupling would need to look like — sustained, feedback-rich, trust-based supervisory relationships — for supervision to function as a genuine improvement lever rather than a compliance ritual. Where loose-coupling theory explains why UAE supervisory policy does not translate uniformly into classroom practice, instructional leadership theory specifies what should be tightly coupled if the system is to produce the outcomes it claims to seek.
2.3 Integrated analytic model
Combining the two traditions, we conceptualize UAE educational supervision as a three-layer system — governance, institutional/school leadership, and classroom/instructional — whose coupling tightness increases (in principle) as it approaches the classroom, moderated by evaluator capacity, workforce diversity, psychological safety, and regulatory coherence. This integrated model, developed inductively from the reviewed literature and presented visually in Figure 1 (Section 8), distinguishes this paper from prior UAE syntheses that treat governance, leadership, and classroom practice as separate, unconnected literatures.
Concretely, the two theories are applied as follows. Section 4 uses loose-coupling theory to read the UAE's MOE-ADEK-KHDA governance structure as a system of formally linked but locally discretionary authorities. Section 5 uses both theories jointly to position the UAE on an international accountability-vs.-trust continuum, drawing on meta-analytic evidence for the claim that inspection design choices generate systematically different improvement mechanisms. Section 6 uses instructional leadership theory to evaluate whether current classroom-level supervisory practices exhibit the tight coupling that theory predicts is necessary for improvement. Section 7 uses loose-coupling theory to reframe the five recurring implementation challenges as manifestations of a single structural coupling problem. Table 4 (Section 4.6) summarizes the UAE-specific evidence base and Table 6 (Section 5.6) presents the international comparative benchmark.
2.4 Operationalizing the framework: observable indicators
To move from metaphor to an analyzable framework, Table 2 specifies, for each construct used in this review, an observable indicator that a future UAE study could code directly from documents, ratings, or interview data. These indicators are used consistently in Sections 4, 6, 7 to support claims about coupling rather than asserting them narratively.
Table 2
| Construct | Definition | Observable indicator(s) in the UAE evidence |
|---|---|---|
| Loose coupling | Elements linked by shared language/goals but not a consistent, predictable operational dependency | Standards vocabulary shared across MOE/ADEK/KHDA documents while implementation protocols and evaluator training differ by authority |
| Tight coupling | Two elements linked strongly, consistently, and predictably | Inter-rater/inter-cycle consistency of observation ratings; a traceable link between specific feedback and subsequent, observable change in practice |
| Vertical coupling | Linkage between hierarchical levels: regulator → school → classroom | Degree to which national/emirate standards language is reproduced, unmodified, in school-level observation instruments and teacher performance plans |
| Horizontal coupling | Linkage across parallel units at the same level (schools within a zone; MOE vs. ADEK vs. KHDA) | Consistency of inspection/evaluation criteria and reporting formats across the three regulatory zones; presence of a shared data dictionary |
| Policy-practice decoupling | Gap between formally adopted policy and everyday enacted practice | Teacher/principal reports of feedback being absent, superficial, or judgmental despite documented completion of the formal observation cycle |
| Regulatory coherence | Extent to which multiple authorities’ requirements are compatible and non-duplicative | Count of overlapping/duplicate reporting requirements across MOE/ADEK/KHDA; existence of a shared definition of effective teaching |
| Responsiveness | Degree to which the system adapts resourcing/practice based on classroom-level evidence | Documented instances of observation/evaluation findings triggering PD resourcing or policy revision, vs. filing without downstream action |
Observable, codable indicators for each coupling construct used in this review.
With these indicators in place, the claim that UAE governance is “loosely coupled” becomes a coded, checkable claim (shared vocabulary plus divergent implementation protocol) rather than an unfalsifiable metaphor, and the claim that classroom practice is “tightly coupled when functioning” is tied to the specific indicator of a traceable feedback-to-practice link, which the UAE evidence (§6.5) shows is inconsistently present.
3 Review methodology
This is a structured narrative Policy and Practice Review, prepared in the Policy and Practice Review format of Frontiers in Education, benchmarked against the field-standard systematic-synthesis approach of , which screened approximately 4,800 studies down to 219. It does not claim the exhaustiveness of a PRISMA-governed systematic review or the protocol pre-registration of a scoping review; the search and synthesis process is reported transparently below so it can be reproduced or extended, and its scope is calibrated explicitly against a systematic-review benchmark in Section 5.4.
3.1 Search strategy and sources
Sources were identified through structured searches of Scopus, ERIC, Google Scholar, and Semantic Scholar, supplemented by direct retrieval from the three UAE regulatory authorities’ official publication portals (MOE, ADEK, KHDA). Search terms combined the concept anchors — supervision, instructional leadership, teacher evaluation, principal evaluation, instructional coaching, school inspection, or classroom observation — with either United Arab Emirates, UAE, or the broader regional term Gulf Cooperation Council/GCC. International theoretical and comparative sources on supervision, coupling theory, instructional leadership, and AI-assisted feedback were retrieved without geographic restriction to establish the theoretical frame in Section 2 and the comparative benchmark in Section 5. Backward-citation harvesting from three anchor meta-analyses and one systematic review of the UAE leadership literature () was used to identify additional international and UAE sources not surfaced by keyword search.
3.2 Inclusion and exclusion criteria
Sources were included if they (a) were peer-reviewed empirical studies, peer-reviewed reviews, or official regulator policy documents; (b) directly addressed supervision, evaluation, instructional leadership, or a directly analogous construct (e.g., instructional coaching, classroom observation reform); and (c) for UAE-specific sources, were published between 2011 and 2025, with priority given to sources from 2020 to 2025. International theoretical and meta-analytic sources were included regardless of publication year when they were the canonical reference for a construct (e.g., ; ). Sources were excluded if they were opinion pieces without an empirical or policy basis, conference abstracts without full text available, or duplicative of a more complete published version of the same study.
3.3 Synthesis procedure
Screening yielded 54 sources retained for full analysis: 15 UAE-specific empirical or policy sources and 39 international theoretical, comparative, meta-analytic, or methodological sources. Sources were thematically coded by hand against the three-layer analytic model described in Section 2.3 (governance, institutional, classroom) and against five recurring implementation themes that emerged inductively during coding: role ambiguity, workforce diversity, evaluator capacity, organizational constraints, and regulatory fragmentation. These themes structure Sections 4, 6, 7. Section 5 additionally benchmarks the UAE system against four contrasting international supervision models, corroborated where possible by meta-analytic evidence. Each of Sections 4, 5, 6, 7 closes with a critical synthesis subsection that explicitly evaluates the quality, consistency, and evidentiary limits of the sources coded to that theme. Table 7 (Section 6.6) summarizes the meta-analytic effect-size evidence used to benchmark UAE findings, so readers can assess directly how UAE claims compare to internationally pooled magnitudes.
3.4 Limitations of the review methodology
As a narrative Policy and Practice Review rather than a systematic review, this synthesis does not report a formal PRISMA flow diagram, inter-rater reliability for coding, or a quality-appraisal instrument scored against every source. Four further limitations bear directly on how the findings and recommendations should be read.
First, publication bias: because searches drew on published, peer-reviewed UAE studies and openly available regulator documents, supervisory practices or outcomes that were attempted but not published — including internal regulator evaluations, unpublished ministry-commissioned reports, and dissertations not indexed in Scopus/ERIC — are systematically underrepresented. The direction of that bias in UAE education research is not neutral: regulator publications skew toward reporting policy-consistent successes.
Second, reliance on policy documents and secondary sources: a substantial share of the UAE governance description in Section 4 rests on regulator-published policy documents (; , ; , ) rather than independent evaluations of those policies’ implementation or effectiveness, so claims about what these frameworks require are more reliable than claims about how well they function in practice.
Third, thin UAE empirical base: as documented in Section 6.5 and Table 4, a small number of overlapping author teams and a heavy reliance on small-sample qualitative case studies constrain the independence and generalizability of the evidence this review synthesizes. Where UAE claims are corroborated by international meta-analytic evidence (Section 6.6, Table 7), we say so explicitly; where they are not, we flag the extrapolation.
Fourth, language and geographic scope: the review is restricted to English-language sources, which likely under-represents Arabic-language work on UAE and GCC supervision and biases the evidence base toward internationally indexed journals.
This review inherits and makes explicit these limitations of the field rather than obscuring them, and Sections 9, 11 flag which recommendations follow directly from reviewed evidence vs. which extend beyond it into professional judgment informed by the comparative analysis in Section 5.
4 Conceptual and policy foundations of supervision in the UAE
4.1 Conceptualizing educational supervision in contemporary leadership scholarship
Educational supervision has been reconceptualized as a central dimension of instructional leadership rather than a peripheral compliance function. Contemporary scholarship positions supervision as a structured process through which leaders influence teaching quality, professional learning, and school improvement (; ). Three conceptual traditions remain especially influential.
Clinical supervision emphasizes structured observation and reflective dialogue, and recent empirical work continues to endorse its formative logic even where implementation is uneven: a 2026 quantitative study of clinical-supervision cycles in public schools found that despite barriers including limited time, insufficient supervisory training, and administrative workload, teachers reported positive effects on instructional planning, classroom management, and reflective practice (). This finding mirrors the older clinical-supervision literature: the practice is durable because it aligns feedback with the specific instructional problem observed, even when institutional support is weak.
Developmental or differentiated supervision recognizes teacher career stages and varied support needs (). Consistent with this developmental tradition, coaching research distinguishes generic or delayed feedback, which produces weak or null effects, from practice-proximal, dialogic feedback tied to standards-aligned observation, which produces larger competence gains (; ).
Instructional leadership models frame supervision as a leadership responsibility directly tied to learning outcomes (; ), a linkage now supported by three independent meta-analytic and systematic-review syntheses (; ; ). These traditions, read through the coupling lens introduced in Section 2, provide the theoretical scaffolding for examining supervisory systems in reform-intensive contexts such as the UAE.
4.2 The UAE reform context and multi-level governance of supervision
The UAE education system is characterized by rapid reform, strong accountability frameworks, and a multi-authority governance structure. The UAE School Inspection Framework () establishes national quality indicators for teaching, learning, leadership, and student outcomes. In Dubai, the Knowledge and Human Development Authority (KHDA) publishes annual inspection findings that publicly categorize school performance and emphasize leadership effectiveness as a driver of improvement (), operationalized through KHDA's own inspection framework document for Dubai private schools (). In Abu Dhabi, the Department of Education and Knowledge (ADEK) formalizes internal self-evaluation, development planning, and continuous monitoring processes that position school leaders as key agents of quality control and improvement ().
These three authorities are not equivalent or uniformly overlapping regulators; they differ in jurisdiction, legal basis, sector coverage, and function. Table 3 makes these differences explicit before the coupling analysis that follows.
Table 3
| Authority | Jurisdiction/sector | Regulatory instrument | Primary function |
|---|---|---|---|
| MOE (Ministry of Education) | Federal; public schools nationwide, plus baseline licensing standards applying across all emirates | UAE School Inspection Framework (); national licensing standards | Sets national quality indicators and teacher/leader licensing requirements |
| ADEK (Abu Dhabi Department of Education and Knowledge) | Abu Dhabi emirate only; public and private schools in Abu Dhabi | Internal self-evaluation and development-planning framework () | Internal, developmentally oriented quality assurance and continuous monitoring |
| KHDA (Knowledge and Human Development Authority) | Dubai emirate only; private schools in Dubai | KHDA Inspection Framework for Dubai private schools (); annual public inspection ratings () | External, publicly reported inspection with a differentiated, publicly consequential rating |
Jurisdiction, regulatory instrument, and primary function of MOE, ADEK, and KHDA.
The three authorities are not interchangeable: MOE sets federal baseline standards, ADEK administers internal self-evaluation within Abu Dhabi, and KHDA administers external, publicly reported inspection within Dubai's private sector — different populations of schools, not repeated measurement of the same population.
The reform intensity is not incidental. Thorne, 2011 study of principal work during the early Abu Dhabi public-private partnership reforms described an environment of intense, urgent national pressure for school reform, in which principals struggled to reconcile their traditional roles with new performance expectations. A decade later, quantitative survey of 113 public-school principals reported substantial adoption of contemporary leadership practices (mean = 4.24 on a 5-point scale), with strategic management the dominant orientation and with statistically significant gender and experience differences in adoption patterns. This trajectory — from urgency-driven reform to formalized leadership standards — is consistent with the fragmented-centralization pattern describes.
In loose-coupling terms, this governance configuration formally links accountability instrument and leadership-development mechanism through shared standards language, while leaving substantial latitude for emirate- and school-level interpretation — the structural source of the fragmentation documented in Section 6. tiered-coupling result specifies why this fragmentation is predictable rather than aberrant: federal policy, district-school relationships, and individual principals each exert partially independent influence on school-level coupling, so a three-authority system is structurally more heterogeneous than a single-regulator system, holding all other design choices constant.
4.3 Professional standards, licensing, and supervision by design
The MOE's Teacher Licensing System (TLS) reflects a national commitment to competency-based regulation, extending supervision beyond episodic evaluation toward ongoing professional validation and development (). UAE-based empirical work supports the importance of this shift: demonstrates how principal evaluation processes in the UAE can either support reflective leadership practice or reinforce bureaucratic compliance, depending on how supervisory routines are enacted, and show that portfolio-based evaluation systems can promote evidence-informed reflection when structured as developmental tools rather than administrative requirements.
Two adjacent UAE studies clarify the supply-side limits on this design. found that public-school principals were dissatisfied with the current three-stage recruitment process (application, interview, probation) and recommended a dedicated Principalship Diploma covering the fundamental core content, training, and requirements for educational leadership. mixed-methods study of transformational leadership in the UAE found that principals believed they were practicing high levels of transformational leadership, while the majority of teachers disagreed — a perception gap that Litz interprets, using Hofstede's cultural frame, as the predictable consequence of layering a Western leadership paradigm onto a differently structured cultural context. Both studies suggest that the professional-standards architecture is more developed than the evaluator- and principal-preparation architecture required to enact it — a coupling gap between design and capacity.
4.4 Fragmentation, coherence, and the policy-practice interface
One of the defining features of the UAE supervisory landscape is regulatory plurality. While MOE frameworks provide national anchors, KHDA and ADEK operate distinct inspection models and reporting systems. Comparative international research suggests that coherence across policy instruments enhances the effectiveness of supervisory systems (); read against loose-coupling theory, this is best understood not as a design flaw to be eliminated but as a structural trade-off between local adaptability and system-wide coherence that UAE policy has not yet explicitly resolved. analysis of the parallel U.S. trajectory — toward fragmented centralization rather than either full tightening or full loosening — offers a productive framing: the UAE is not choosing between coupling and decoupling but between competing patterns of selective tightening.
4.5 Critical synthesis
The sources reviewed in this section converge on a shared description of the UAE's governance architecture but stop short of explaining, in theoretical terms, why that architecture produces the implementation variability documented later in the review (, ) studies, while methodologically careful, are single- or small-multi-case qualitative designs situated in Abu Dhabi public schools; their findings on principal evaluation are not yet tested against Dubai's private-sector KHDA context or against MOE federal schools, leaving an untested assumption of transferability across the three regulatory zones. Similarly, and are self-published regulator documents rather than independently peer-reviewed evaluations of their own policies, a source-independence limitation this review treats as a constraint on the strength of claims about policy effectiveness. The transformational-leadership perception gap reports — principals rating themselves high while teachers rate them substantially lower — is also a warning signal: self-reported adoption of leadership practices () may systematically overstate what teachers experience.
4.6 Summary of the UAE-specific evidence base
Table 4 addresses the reviewer request for a comparative overview of the reviewed studies’ design, scope, and findings, summarizing the fifteen UAE-specific sources this review draws on.
Table 4
| # | Study (Year) | Design | Scope/Sample | Regulatory zone | Key finding |
|---|---|---|---|---|---|
| 1 | Thorne, 2011 | Qualitative case study | 1 principal, reform period | Abu Dhabi (PPP schools) | Reform pressure reshaped principal role faster than preparation infrastructure could keep pace |
| 2 | Mixed methods | 130 survey; 4 principals, 12 teachers | UAE (mixed) | Principals rate own transformational leadership high; teachers disagree — cultural-fit gap | |
| 3 | Scoping review | UAE leadership literature | UAE | Reform-driven, multicultural environment requires context-sensitive leadership | |
| 4 | Qualitative multi-case | Abu Dhabi public schools | Abu Dhabi | Principal evaluation supports reflection or bureaucratic compliance depending on enactment | |
| 5 | Qualitative | UAE principals | UAE | Feedback often absent, superficial, or judgmental | |
| 6 | Qualitative case study | UAE portfolio evaluation | UAE | Portfolios support evidence-informed reflection when developmental | |
| 7 | Narrative inquiry | Remote-school language teachers | UAE (remote) | Coaching-oriented supervision supports PD in dispersed settings | |
| 8 | Qualitative multi-case | UAE teachers | UAE | Teachers perceive evaluation as compliance-oriented; limited reflection | |
| 9 | Qualitative | UAE school leaders | UAE | Regulatory pressure and workload limit developmental supervision | |
| 10 | Scoping review | UAE PD literature 2018-23 | UAE | Supervisory evidence most useful when feeding coherent PD | |
| 11 | Qualitative | Public-school principals | MOE (federal) | Principals dissatisfied with recruitment; call for principalship diploma | |
| 12 | Quantitative survey | 113 principals | UAE public | High self-reported adoption of contemporary leadership (mean 4.24); strategic management dominant | |
| 13 | Systematic review | UAE instructional leadership | UAE | Thematic synthesis of instructional leadership in the UAE | |
| 14 | Regulator policy document | Abu Dhabi private schools | ADEK | Self-evaluation, development planning, continuous monitoring | |
| 15 | ; | Regulator policy + inspection findings | Dubai private schools | KHDA | Publicly reported, differentiated inspection outcomes |
UAE-specific evidence base on educational supervision and leadership.
The patterns visible here — predominance of small-sample qualitative designs, concentration in Abu Dhabi public schools, overlapping author teams, and reliance on regulator publications — are the empirical basis for the critical synthesis in Section 4.5 and are referenced again in Sections 3.4, 6.5, and 7.6.
5 International comparative perspectives on educational supervision
This section benchmarks the UAE's supervisory system against four contrasting international models, selected to represent distinct points on the accountability-vs.-trust and centralization-vs.-professionalization continua that Section 2 identifies as the theoretical crux of supervision design. Each subsection is corroborated, where possible, by meta-analytic or systematic-review evidence rather than a single primary source.
5.1 Rationale for comparator selection
The four comparators were selected not as a convenience sample of well-known systems but because each anchors a distinct, empirically documented point on the two continua identified in Section 2 as the theoretical crux of supervision design: accountability stakes (low-to-high) and developmental-infrastructure density (low-to-high). Comparison is meaningful only on these two structural dimensions and the mechanisms tied to them — inspection design, trust logic, career-track architecture, and evidence-synthesis rigor — not on curriculum content, assessment systems, or broader education outcomes, which differ too much across these systems to support direct transfer.
These four systems are benchmarked on accountability design and developmental infrastructure only; differences in workforce composition, teacher-education pipelines, and curricular plurality mean that a design feature effective in one system cannot be assumed transferable to the UAE without local piloting and evaluation (Section 9). Table 5 presents the rationale for comparator selection, anchored to the accountability-stakes and developmental-infrastructure-density continua.
Table 5
| Comparator | Why selected | Dimension it anchors | Contextual differences limiting transfer |
|---|---|---|---|
| Finland | Most-cited empirical counter-example of accountability substituted by professional trust, with a documented teacher-training pipeline underpinning that trust () | Low accountability stakes/trust-based development | Small, linguistically homogeneous population; universal, selective master's-level teacher education; no equivalent to the UAE's expatriate-majority workforce () |
| England (Ofsted) | Clearest documented case of high-stakes, publicly reported inspection with measured workforce costs (); closest structural analogue to KHDA's public-rating model (§4.2) | High accountability stakes/low developmental infrastructure | Unitary national system, domestically trained mono-cultural workforce, no multi-curriculum private-school landscape |
| Singapore | Clearest documented case of centralization paired with an explicit competency-track architecture converting appraisal into career development (); small, high-growth state comparable to the UAE in scale/reform tempo | High centralization/high developmental infrastructure | More ethnically homogeneous workforce; single national curriculum vs. the UAE's multiple co-existing curricula |
| United States (post-RTTT) | Substantive comparator (large-scale, multi-jurisdiction teacher-evaluation reform) and the field's methodological benchmark for systematic evidence synthesis () | Federated, multi-jurisdiction design; methodological rigor benchmark | District-level (not emirate-level) devolution; unionized, domestically credentialed workforce; decades of quantitative data the UAE lacks |
Rationale for comparator selection, anchored to the accountability-stakes and developmental-infrastructure-density continua identified in Section 2.
5.2 High-stakes inspection systems: Europe and England
A comparative study of six European inspection systems (the Netherlands, England, Sweden, Ireland, the Austrian province of Styria, and the Czech Republic) found that inspection models differ systematically in visit scheduling, the balance of process vs. output standards, and the consequences attached to findings, and that these design choices generate different mechanisms of school improvement rather than a single universal effect (). Recent large-scale evidence from England reinforces the darker side of this pattern: the Beyond Ofsted survey of teachers and school leaders found that 76% of respondents believed Ofsted had a negative effect on retention, with 30% considering leaving the profession as a result of their most recent inspection, and many respondents described the inspection regime as “toxic and brutal” ().
This evidence is directly relevant to the UAE. KHDA's publicly reported, differentiated inspection model () most closely resembles the higher-stakes, publicly consequential inspection regimes in the European comparison, whereas ADEK's self-evaluation-anchored model () sits closer to the lower-stakes, developmentally oriented end of the same continuum. The English evidence suggests these design choices are not neutral: publicly consequential inspection can generate targeted improvement but also produces measurable workforce costs that the UAE literature reviewed in Section 6 has not yet tested empirically within its own jurisdictions.
5.3 Trust-based, low-stakes supervision: Finland
Finland represents the opposite pole of the accountability-trust continuum: its system deliberately limits external testing and inspection in favor of placing responsibility and trust in highly selected, master's-level-trained teachers, with school- and district-level leadership treated as an extension of that professional trust rather than as external oversight (). Read against the UAE evidence in Sections 6 and 7, the contrast is instructive: the psychological-safety and trust deficits UAE teachers and principals report (; ) are precisely the conditions the Finnish model is designed to avoid by minimizing external accountability pressure in the first place. The English retention evidence and the Finnish trust-based counter-example jointly suggest that the accountability-development tension is a predictable systemic property of any inspection regime that is not paired with compensating developmental infrastructure, not a feature specific to any one country.
This does not imply the UAE should adopt a Finnish-style low-stakes model — the two systems differ enormously in workforce composition, curricular diversity, and reform tempo (see Section 5.5) — but it does clarify that the UAE's trust deficit is a structural consequence of pairing high external accountability with underdeveloped internal coaching capacity, not an isolated implementation failure.
5.4 Competency-based, career-track supervision: Singapore
Singapore offers a third model: a centralized but developmentally structured system in which supervision is embedded in explicit competency frameworks and career tracks (Teaching Track, Leadership Track, Senior Specialist Track), with performance management designed to identify and cultivate leadership capacity progressively rather than through a single evaluative event (). This model demonstrates that centralization and developmental supervision are not inherently in tension, and it offers a plausible benchmark for GCC systems considering how to convert accountability data into structured developmental pathways. UAE finding that principals themselves called for a dedicated Principalship Diploma suggests domestic appetite for a more structured career pathway. The absence of an equivalent empirically evaluated career-track architecture in the UAE evidence base is a concrete, actionable gap rather than a generic critique.
5.5 Evidence synthesis as a methodological benchmark: the United States
Beyond substantive models, international scholarship also offers a methodological benchmark relevant to Section 3's transparency aims. synthesis of principal effects screened approximately 4,800 studies down to 219 meeting explicit relevance and rigor criteria. Complementing the Grissom synthesis, the meta-analytic evidence base has expanded rapidly: meta-meta-analysis of twelve prior meta-analyses reports d = 0.34 ; extend this with a 42-study multivariate meta-analysis reporting 0.22–0.25 SD depending on modeling choice; and systematic review of 51 studies reports the direct-effect magnitudes noted in Section 1. The comparison is included here precisely so that reviewers and readers can calibrate the confidence warranted by this review's UAE-specific claims against a field-standard example of what fuller systematic rigor looks like.
5.6 GCC-regional positioning: the workforce-composition constraint
A dimension not adequately addressed in prior UAE supervision reviews is the regional workforce structure document that expatriate Arab teachers make up a substantial share of the government-school teaching workforce in the UAE and Qatar, with the proportion varying markedly by school gender-stream. A systematic review of novice teachers’ perceptions of Gulf teacher-education programs identifies recurring weaknesses — an impactful but insufficient practicum, a theory-practice gap, and non-culturally responsive curricula — across multiple GCC countries (). Together these findings suggest that the developmental supply-side of UAE supervision — the pipeline of teachers whom supervision is meant to develop — inherits regional constraints that any policy reform must reckon with. A Finnish-style trust-based supervision presupposes a workforce trained through a stable, culturally embedded teacher-education pipeline; the workforce composition documented for the region places different, not lower, demands on supervisory design.
5.7 Critical synthesis and positioning
Placed side by side, the four comparators show that the UAE's supervisory system is not a novel configuration but an unresolved hybrid of existing models: structurally closer to the high-stakes, publicly reported inspection paradigm of Europe and England than to Finland's trust-based model, yet without Singapore's explicit competency-track infrastructure for converting accountability data into developmental pathways. Table 6 summarizes the four comparators against six analytic dimensions relevant to the UAE case.
Table 6
| Dimension | England (Ofsted) | Finland | Singapore | US (post-RTTT) | UAE (current) |
|---|---|---|---|---|---|
| Accountability stakes | Very high; public ratings | Very low; no external inspection | Moderate; internal to career track | High; district-level | Mixed (KHDA high, ADEK moderate, MOE moderate) |
| Trust logic | Low; audit-based | High; profession-based | Structural; competency-track-based | Low; audit + value-added | Emerging; not systematically designed |
| Career-track architecture | Weak | Implicit via profession | Explicit, three tracks | Weak | Absent |
| Documented workforce cost | 76% negative retention effect | Not applicable | Low reported strain | Widget-effect ratings inflation | Under-documented |
| Meta-analytic corroboration | Inspection-effect literature mixed | Not applicable | Career-track design not meta-analyzed | Coaching d ≈ 0.49; principal d ≈ 0.34 | UAE-specific meta-analytic base absent |
| Coupling posture | Tight at classroom, tight at governance | Loose at governance, tight-by-trust at classroom | Tight throughout, developmentally sequenced | Fragmented centralization | Loosely coupled at governance, uneven at classroom |
Comparative supervision-model benchmark against the UAE, corresponding to the analysis in Section 5.6.
This positioning is itself a critical, not merely descriptive, claim: it argues that the accountability-development tension documented throughout Sections 6 and 7 is not an implementation defect to be trained away but a structural consequence of importing a high-stakes inspection logic without the compensating developmental infrastructure that either the Finnish (trust) or Singaporean (career-track) models supply. None of the UAE-specific sources reviewed in this manuscript makes this comparative argument explicitly, which is the basis for treating it as an original contribution of this review (see Section 10).
6 Current supervisory practices in UAE schools
6.1 Instructional supervision approaches
Classroom observation remains the most visible supervisory practice across UAE school systems and curricula, used for both formative feedback and summative judgements. In Dubai, observation evidence is embedded in inspection processes, where inspectors conduct classroom visits and triangulate evidence across outcomes, teaching quality, and leadership (). In Abu Dhabi, ADEK's quality assurance policy similarly emphasizes internal monitoring of teaching quality standards and evidence-based improvement planning (). However, the quality of supervision depends heavily on the feedback and coaching that follow observations — an issue underscored in UAE studies where feedback is reported as inconsistent, episodic, or overly compliance-focused (), echoing earlier UAE evidence that principals themselves report absent, superficial, or judgmental supervisory feedback that limits professional learning dialogue ().
International evidence corroborates the mechanism the UAE studies suggest. showed in the Cincinnati observation study that trained, external observers using an extensive set of standards can reliably identify effective teachers and teaching practices, but found in a randomized Chicago pilot that observation reform improved student reading performance only in low-poverty, high-achieving schools with strong central-office support, and that the same program produced no such effect when central-office support waned in the following year. Observation quality is therefore a necessary but not sufficient condition for improvement, and the sufficient conditions concern support infrastructure — a pattern directly relevant to the UAE regulatory-fragmentation problem.
6.2 Supervisor roles and responsibilities
Supervision in UAE schools is enacted through layered leadership roles, typically including principals, vice principals, heads of department, subject coordinators, and designated instructional leaders, operating alongside regulator-related roles such as inspectors and cluster supervisors. Principals experience supervision both as a responsibility (supervising teachers) and as a target (being supervised through evaluation); UAE studies highlight how these dual pressures can intensify workload and shape supervisory priorities ().
6.3 Collaborative and teacher-centered supervisory practices
Professional learning communities, peer collaboration structures, and mentoring initiatives are increasingly visible across UAE schools. UAE research on teacher evaluation suggests that when supervision is experienced primarily as accountability, teachers may adopt compliance behaviors rather than engage in authentic professional inquiry (); conversely, coaching-oriented supervision appears to support genuine professional development, including in remote and geographically dispersed settings (). The mechanism connecting coaching to instructional change has been rigorously specified in the international literature: meta-analysis of 60 causal studies reports coaching effects of 0.49 SD on instruction and 0.18 SD on student achievement, and process-level research shows that coach-teacher relationship quality predicts teacher engagement and reflection, while coaching-strategy quality predicts overall classroom instructional quality ().
6.4 Tools and technologies shaping current practice
Supervision is increasingly mediated through digital systems: lesson observation templates, teacher performance documentation, electronic portfolios, and school data dashboards. These tools can enhance consistency and transparency but may also encourage performative documentation if feedback and coaching capacity are weak (; ). A related scoping review of UAE professional development literature likewise finds that supervisory evidence is most useful when it feeds coherent, needs-based PD planning rather than isolated reporting requirements ().
6.5 Critical synthesis
Across the practice-level literature, a consistent finding is that observation and documentation infrastructure is well developed across all three regulatory zones, while the feedback and coaching layer that instructional leadership theory identifies as the active mechanism for improving teaching (Section 2.2) is the least consistently implemented and the least rigorously studied element. Notably, the strongest claims about coaching-oriented supervision's benefits (; ) originate from an overlapping set of UAEU-affiliated researchers using qualitative case designs; this pattern strengthens internal coherence of the findings but weakens claims to independent replication. The practice literature also over-represents private and semi-private schooling; MOE federal public-school supervisory practice is comparatively under-documented, a gap future UAE research should prioritize before the practice-level claims in this section are generalized system-wide.
The UAE evidence is also almost entirely qualitative and perception-based. No UAE study identified in this review quantitatively estimates the effect of supervisory practice on teacher instruction or student learning at anything approaching the sample sizes and design rigor of the international meta-analyses summarized in Table 7.
6.6 International contextual benchmarking for UAE practice claims
Table 7 summarizes effect magnitudes at a glance; Table 8 provides the fuller methodological detail needed to judge how much weight each benchmark can bear, and states explicitly what each source is relevant to the UAE claims this review makes.
Table 7
| Construct | Source (Year) | k | Effect | Interpretation |
|---|---|---|---|---|
| Principal leadership → student achievement | 12 meta-analyses | d = 0.34 | Small-to-moderate; robust across syntheses | |
| Principal leadership → student achievement (US, 2000-2020) | 42 | 0.22–0.25 SD | Contingent on direct/indirect modeling | |
| Principal behaviors → student achievement | 51 | 0.08–0.16 SD | Direct effects small; teacher-outcome effects larger | |
| Principal behaviors → instructional practice | 51 | 0.35 SD | Moderate | |
| Principal behaviors → teacher well-being | 51 | 0.34–0.38 SD | Moderate | |
| Teacher coaching → instruction | 60 (causal) | 0.49 SD | Moderate-to-large; scales down in effectiveness trials | |
| Teacher coaching → student achievement | 60 (causal) | 0.18 SD | Small-to-moderate | |
| Pre-service coaching/mentoring/supervision → instructional skills | 12 | d = 0.41 | Small; cognitive modeling moderates strongly (d = 0.90) | |
| Principal transformational leadership → teacher job satisfaction | 10 | r = 0.50 | Moderate | |
| Academic press → student achievement | 79 (30 years) | Large | Robust and leadership-malleable |
International meta-analytic effect-size benchmarks for constructs asserted in the UAE supervision literature.
Table 8
| Source | Population | Intervention | Comparison | Outcome | Design | Effect metric | Heterogeneity | Risk of bias | Relevance to UAE claim |
|---|---|---|---|---|---|---|---|---|---|
| K-12 schools, 12 meta-analyses pooled | Principal leadership practices (general) | Lower- vs. higher-leadership schools | Student achievement | Meta-meta-analysis | Cohen's d = 0.34 | Not formally quantified; narratively moderate | Depends on quality of 12 underlying meta-analyses; not independently appraised here | General-leadership benchmark only — not supervision-specific (§1.1); contextualizes UAE principal-leadership claims | |
| K-12 teachers, 60 causal studies | Teacher coaching (observation + feedback) | No-coaching or business-as-usual control | Instruction (0.49 SD) and achievement (0.18 SD) | Meta-analysis of causal (RCT/quasi-experimental) studies | SD (Hedges’ g) | Substantial; scaled-up effectiveness trials show markedly smaller effects than efficacy trials | Not independently appraised; authors note publication-bias checks | Direct benchmark for UAE supervision/coaching claims (§6) | |
| Pre-service teachers, 12 studies | Coaching, mentoring, and supervision (pre-service) | No/minimal supervision control | Instructional skill acquisition | Meta-analysis | d = 0.41 overall; d = 0.90 for cognitive-modeling subgroup | Moderate; cognitive modeling identified as a strong moderator | Not independently appraised | Benchmark for evaluator-capacity recommendation in §7.3/§9.5 | |
| K-12 schools, 51 studies | Principal behaviors (instructional + organizational) | Comparison schools/leaders | Student achievement, instructional practice, teacher well-being, school health | Systematic review + meta-analysis | SD across four outcome families (0.08–0.81 SD) | High; effect size varies by outcome family | Not independently appraised | Distinguishes leadership from supervision effects (§1.1); UAE has no equivalent multi-outcome study | |
| ; | School inspectorates and teachers, England + 5 European systems | External inspection (design variation) | Cross-system comparison | Improvement mechanism; teacher retention/well-being | Comparative case study (Ehren); large-scale survey (Perryman) | Descriptive/qualitative; % negative-impact reports (Perryman) | Not applicable (non-meta-analytic) | Single-country survey (Perryman) not independently appraised for representativeness | Direct comparator for KHDA's inspection model (§5.1) |
| Singapore teaching workforce | Competency-track career architecture | Pre-reform/non-tracked systems | Career progression, appraisal-to-development conversion | Policy/system-design case study | Not meta-analytic; descriptive | Not applicable | Single-system case study; not independently appraised | Comparator for the conditional Singapore-track recommendation (§9.6) |
International contextual benchmarking detail: population, intervention, comparison, outcome, design, effect-size metric, heterogeneity, risk-of-bias status, and relevance to the specific UAE claim each source is used to benchmark.
“International contextual benchmarking” denotes the use of these sources as contextual reference points for interpreting UAE-specific findings, not as meta-analytic corroboration of a pooled UAE effect — no UAE studies are included in the pooled estimates themselves.
Two implications follow. First, the UAE literature's qualitative claim that coaching-oriented supervision produces reflection and instructional improvement is directionally consistent with the strongest available international meta-analytic evidence, but the UAE literature does not yet establish comparable effect magnitudes for its own context. Second, the meta-analytic evidence itself is heterogeneous — direct effects are small, indirect and teacher-outcome effects are larger, and effect sizes in scaled-up effectiveness trials are meaningfully smaller than in tight efficacy trials — so UAE policy design should treat the meta-analytic magnitudes as upper bounds on what a well-implemented reform can plausibly achieve, not as guaranteed returns.
7 Challenges facing educational supervision in the UAE
7.1 Role ambiguity and the accountability-development tension
A persistent challenge in supervisory systems internationally concerns the dual purpose of supervision: accountability and professional growth. Contemporary scholarship emphasizes that evaluation systems designed primarily for accountability can undermine developmental supervision if not carefully balanced (; ). This is not merely a rhetorical concern: demonstrate empirically, in a multi-case study of eight high-functioning elementary principals, that supervision and evaluation constitute a wicked problem — tensions to be acknowledged and worked with, not resolved through a design fix. found that UAE teachers often perceived evaluation processes as compliance-oriented, with limited opportunities for reflective dialogue — a finding the international literature would predict for any high-stakes system without a compensating developmental layer.
A further empirical warning comes from analysis of teacher evaluation reforms across two dozen U.S. states adopting major post-Race-to-the-Top reforms: despite considerable design investment, unsatisfactory ratings remained below 1% in the vast majority of states, and evaluators in a surveyed urban district privately perceived far more teachers to be below proficient than they formally rated as such. This persistence of undifferentiated ratings even after ostensibly reformed systems are in place is directly relevant to the UAE, where the perception among teachers of compliance-oriented evaluation would be predicted to worsen if principals face implicit pressure to avoid negative ratings.
7.2 Workforce diversity and cultural complexity
The UAE education workforce is among the most internationally diverse in the world. document that expatriate Arab teachers historically constituted a substantial share of the UAE government-school teaching workforce, and systematic review of GCC novice-teacher perceptions identified consistent theory-practice and cultural-responsiveness gaps in teacher-education programs. emphasize that UAE school leadership operates within a reform-driven, multicultural environment requiring context-sensitive leadership strategies; supervisors must navigate differences in instructional traditions, professional norms, and communication expectations while maintaining alignment with standardized evaluation criteria finding of a principal-teacher perception gap on transformational leadership is best read against this workforce-composition backdrop: shared vocabulary does not guarantee shared interpretation in a culturally heterogeneous workforce.
7.3 Evaluator capacity and supervisory expertise
Supervisory impact is closely linked to evaluator expertise (; ). reported that principal evaluation processes depended heavily on supervisors’ interpretive competence and professional judgment; where supervisors lacked training in instructional coaching or subject-specific pedagogy, feedback tended to emphasize procedural compliance rather than pedagogical refinement. This finding is consistent with the international coaching literature's identification of cognitive modeling — the supervisor making pedagogical processes explicit and demonstrable — as the largest moderator of coaching effect size (d = 0.90 );. It is also consistent with the process-level finding that coach-teacher relationship quality and coaching-strategy quality jointly predict engagement and instructional improvement (). In coupling terms, evaluator capacity is the practical mechanism by which loose regulatory-level coupling either survives contact with the classroom or degrades into performative documentation.
7.4 Organizational constraints and administrative load
documented stressors among UAE school leaders, including regulatory pressure and workload intensification, which can limit the capacity of leaders to conduct sustained developmental supervision. The English retention evidence makes the mechanism explicit: high-stakes inspection regimes generate workload and psychological costs that at scale reduce the workforce stability required for developmental supervision to accumulate ().
7.5 Teacher perception, psychological safety, and trust
When teachers perceive evaluation systems as punitive or unreliable, they may engage in performative compliance rather than authentic professional learning (; ). In the UAE, qualitative evidence suggests that teachers sometimes question the transparency and consistency of evaluation processes (). demonstration that evaluators privately perceived far more teachers as below proficient than their formal ratings reflected is a warning that the ratings system itself is a social act embedded in the trust environment; when the trust environment is thin, the ratings do not accurately reflect professional judgment.
7.6 Critical synthesis
Read collectively and against the theoretical frame in Section 2, the five challenges above are not independent problems but symptoms of the same underlying structural condition: a governance layer that is loosely coupled to classroom practice by design, without an explicit UAE policy statement of how tight that coupling should be at each level. This reframing is itself an original analytical move of this review, distinguishing it from prior UAE literature (e.g., ; ) that documents these challenges as a list of parallel pressures rather than as manifestations of a single coupling problem.
We also note, as a limitation the field should address, that the UAE challenge literature is almost entirely qualitative and perception-based; no UAE study identified in this review quantitatively links evaluator capacity, feedback quality, or psychological safety to measured instructional improvement or student outcomes. The causal strength of the accountability-development tension described here therefore rests on inference from the international literature — on evaluator behavior, on coaching effects, on cognitive modeling, and on inspection-driven retention costs — rather than a UAE-validated finding. This is a research gap the field should prioritize.
8 Conceptual model: a visual synthesis
Figure 1 translates Sections 2, 4, 6, and 7 into a single conceptual model, informed by the comparative positioning developed in Section 5 and the meta-analytic benchmarks summarized in Table 7. It depicts UAE educational supervision as three layers of decreasing coupling looseness — governance, institutional/school leadership, and classroom/instructional — connected by a downward accountability flow and an upward feedback/evidence flow, and moderated by four cross-cutting factors: evaluator capacity, workforce and curricular diversity, psychological safety and trust, and regulatory fragmentation.
Figure 2 complements the conceptual model by positioning the UAE against the four international comparators (Section 5) on two dimensions: accountability stakes and developmental-infrastructure density, using the comparative evidence base in Table 6 as calibration.
Figure 2
9 Future innovations in UAE educational supervision and leadership
To keep these recommendations aligned with the evidence that supports them, each subsection below opens with an explicit evidence-basis statement distinguishing what is directly supported by cited findings from what extends beyond them into professional judgment informed by the comparative analysis in Section 5 and the meta-analytic benchmarks in Table 7.
9.1 Coaching-oriented supervision
Evidence basis: strong internationally, moderate for the UAE. International meta-analytic evidence links coaching-based supervision to instructional gains of 0.49 SD and student-achievement gains of 0.18 SD, though scaled-up effectiveness trials show smaller effects (). Pre-service coaching effects are smaller (d = 0.41) but strongly moderated by supervisor cognitive modeling (d = 0.90; ). UAE qualitative evidence independently shows that developmental, coaching-style feedback is associated with more reflective practice than compliance-oriented feedback (; ). In coupling terms, coaching-oriented supervision is best understood as a deliberate policy choice to tighten coupling at the classroom layer while keeping evaluative and developmental functions structurally separated ().
Operationally, the international evidence suggests three design principles UAE policymakers should treat as prerequisites rather than aspirations: (i) invest in supervisor cognitive modeling as the largest moderator of effect size; (ii) protect coach-teacher relationship quality, which predicts teacher engagement and reflection; and (iii) expect effects at scale to be smaller than pilot effects, so reform business cases should be built around realistic magnitudes.
9.2 AI-assisted instructional feedback
Evidence basis: emerging internationally, not yet UAE-tested. A single randomized controlled trial found that automated feedback systems can improve teachers’ uptake of students’ ideas in an online-course setting outside the UAE (). A structured integrative review of 37 empirical studies of AI-automated feedback in higher education (2014–2024) concludes that potential benefits — customization, immediacy, engagement — are persistently compromised by concerns regarding algorithmic bias, data privacy, the deterioration of teacher-student relationships, and inadequate professional growth, and that the current evidence base is methodologically deficient, largely short-term or subjective, with few longitudinal or controlled comparisons (). More recent systems research () explores multimodal AI producing temporally precise, best-practice-aligned feedback from classroom video recordings, though this is an early-stage engineering preprint rather than a validated educational evaluation.
No UAE study in this review's evidence base has evaluated AI-assisted supervision. The recommendation that UAE supervisors could use such tools is therefore an extrapolation from non-UAE experimental and design evidence, not a UAE-validated finding, and should be read accordingly: AI functioning as a decision-support mechanism rather than an autonomous evaluative authority, with explicit governance for surveillance, bias, and data-privacy risk.
9.2.1 Ethical governance of AI-assisted supervision
Because no UAE study has evaluated AI-assisted supervision, any deployment must be governed as a carefully evaluated pilot rather than scaled immediately. Nine ethical dimensions require explicit governance before any pilot proceeds; Table 9 specifies the risk and a minimum safeguard for each.
Table 9
| Dimension | Risk if ungoverned | Minimum safeguard for a pilot |
|---|---|---|
| Privacy | Classroom audio/video capture exposes students and teachers beyond the evaluation purpose | Data minimization; capture limited to the observation window; no incidental student-identifying data retained |
| Consent | Teachers evaluated by a system they did not agree to be assessed by | Documented, revocable teacher consent prior to any AI-assisted observation cycle; opt-out without penalty during the pilot |
| Surveillance | Continuous or covert monitoring shifts supervision from developmental support to control | AI tools scoped to scheduled, announced observation windows only — not continuous or unannounced monitoring |
| Data ownership | Ambiguity over who controls recordings/analytics (vendor, regulator, school, teacher) | Written data-ownership and retention agreement naming the school/regulator as owner before any vendor contract |
| Algorithmic bias | Models trained on non-UAE, non-multilingual classrooms misjudge UAE teaching contexts | Bias audit against a UAE-representative sample before deployment; documented performance by language and curriculum |
| Linguistic validity | Feedback models calibrated on English-only speech may misclassify Arabic or bilingual instruction | Validation of the tool's accuracy specifically on Arabic-medium and bilingual classroom speech before use in those settings |
| Data security | Classroom recordings are high-sensitivity data attractive to breach | Encrypted storage, access logging, and a defined retention/deletion schedule aligned to UAE data-protection requirements |
| Appeals | A teacher has no recourse if an AI-generated rating is contested | A human-adjudicated appeals process for any AI-informed rating, with the AI output treated as advisory, not final |
| Human oversight | Automated output is accepted as final without professional judgment | A qualified human evaluator reviews and can override every AI-generated observation before it affects a teacher's record |
| Psychological safety | Awareness of algorithmic monitoring suppresses risk-taking and honest reflection in coaching conversations | Pilot framed and communicated as formative and low-stakes, with explicit protection from AI output being used in high-stakes personnel decisions during the pilot phase |
Ethical dimensions requiring explicit governance before any UAE pilot of AI-assisted supervision.
These safeguards are the minimum condition for treating AI-assisted supervision as a carefully evaluated pilot. Absent all ten, deployment should not proceed beyond a small, consented, appeal-enabled pilot with explicit human oversight — immediate scaling is not supported by the evidence base reviewed here (§9.2).
9.3 Digital evidence systems and data-informed improvement
Evidence basis: descriptive/directional. UAE policy documents describe digital observation templates, portfolios, and dashboards as already in use (), and international research on evaluation reform argues data is most valuable when used for learning rather than reporting (). No UAE study in this review's evidence base directly tests whether digital systems improve instructional outcomes. Innovation here lies not in digitization itself but in ensuring these systems feed reflective dialogue and instructional decision-making rather than becoming compliance archives. finding about undifferentiated ratings is the concrete warning: digitizing a rating process that already produces undifferentiated results merely digitizes the undifferentiation.
9.4 System coherence and regulatory alignment
Evidence basis: moderate, cross-nationally supported but not UAE-tested. Comparative international research links policy coherence across governance instruments to more effective supervisory systems (), and the tiered-coupling literature specifies why fragmented centralization is the predictable steady state absent deliberate design choices. The UAE-specific structural description in this review (; Section 4) documents the regulatory fragmentation such coherence would address. No UAE study has directly tested whether tighter MOE-ADEK-KHDA coherence improves outcomes. Strengthening coherence between MOE inspection frameworks, KHDA reporting structures, and ADEK quality-assurance policies could reduce duplication and reinforce common definitions of effective teaching and leadership — in coupling terms, a deliberate, selectively tighter coupling at the governance layer, applied to shared definitions and data standards rather than to every operational detail.
9.5 Ethical governance and capacity building
Evidence basis: judgment call, logically derived rather than directly evidenced. This recommendation follows from the evaluator-capacity gaps documented in Section 7.3, the cognitive-modeling moderator from the pre-service coaching meta-analysis, and the AI-governance caveats in Section 9.2. It is included because the reviewed evidence base identifies the underlying capacity gap clearly even though no study has evaluated a specific UAE capacity-building intervention. As supervision becomes increasingly data-rich and technologically mediated, evaluator preparation must evolve to include data interpretation, coaching strategy, and ethical technology use.
9.6 A Singapore-style career-track layer for the UAE (conditional recommendation)
Evidence basis: comparative and conditional. document Singapore’s explicit competency-track architecture, and UAE finding that principals themselves called for a dedicated Principalship Diploma suggests domestic appetite for a structured career pathway. We treat this as a conditional recommendation because no UAE trial of such a career-track layer has been evaluated; the recommendation is therefore that it be piloted with an explicit evaluation design rather than adopted at scale on comparative-inference grounds alone.
9.7 Policy recommendation matrix
Table 10 consolidates the six recommendations made across Section 9 (§9.1–§9.6) into a single matrix, specifying for each the responsible authority, implementation level, required resources, expected outcome, evaluation indicator, principal risk, and strength of supporting evidence.
Table 10
| Recommendation | Responsible authority | Level | Resources required | Expected outcome | Evaluation indicator | Principal risk | Evidence strength |
|---|---|---|---|---|---|---|---|
| Coaching-oriented supervision (§9.1) | MOE/ADEK/KHDA jointly | School + regulator | Coach training; protected coaching time | Improved instructional quality via sustained coaching relationships | Coach-teacher relationship quality; observed instructional-quality change | Coaching reduced to another compliance cycle without protected time | Strong (international meta-analytic; UAE-untested) |
| AI-assisted instructional feedback, piloted only (§9.2) | School, with regulator oversight | School (pilot) | Vendor tool; consent process; human-review capacity (Table 9) | Faster, more consistent formative feedback | Pilot completion against the Table 9 safeguards; teacher consent/opt-out rates | Premature scaling without governance (Table 9) | Weak/emerging (non-UAE evidence only) |
| Digital evidence systems for reflective use (§9.3) | ADEK/KHDA | Regulator + school | Dashboard/portfolio infrastructure (largely already in use) | Data used for instructional learning, not just compliance reporting | Proportion of digital records linked to a documented follow-up action | Digitizing an undifferentiated rating system merely digitizes the undifferentiation | Weak (UAE-descriptive only; no outcome test) |
| System coherence and regulatory alignment (§9.4) | MOE, ADEK, and KHDA jointly | Governance | Cross-authority coordination mechanism; shared data dictionary | Reduced duplication; common definition of effective teaching/leadership | Count of harmonized reporting requirements; existence of a shared definition of effective teaching | Coherence effort adds a fourth layer of process instead of reducing the existing three | Moderate (cross-national evidence; UAE-untested) |
| Ethical governance and evaluator capacity building (§9.5) | MOE/ADEK/KHDA | Evaluator workforce | Training curriculum covering data interpretation, coaching strategy, ethical technology use | Evaluators equipped for data-rich, technology-mediated supervision | Evaluator training completion; cognitive-modeling skill assessment | Capacity building treated as a one-time course rather than sustained practice | Judgment call (logically derived; not directly evidenced) |
| Singapore-style career-track layer, conditional pilot only (§9.6) | MOE, with emirate-level regulators | System (pilot) | New career-track architecture; principalship diploma infrastructure | Structured career pathway converting appraisal into development | Pilot cohort retention and progression; principal-reported role clarity | Adoption at scale without a UAE evaluation of transferability | Comparative and conditional (Singapore evidence; no UAE trial) |
Policy recommendation matrix: responsible authority, implementation level, resources, expected outcomes, evaluation indicators, principal risks, and evidence strength for each recommendation in Section 9.
10 Original contribution and positioning
This review's contribution should be read against four bodies of prior work it builds on and departs from.
First, relative to UAE-specific empirical studies (e.g., ; ; ), which each examine a single supervisory mechanism in depth, this review's contribution is synthetic and theoretical rather than empirical: it does not generate new primary data but organizes existing findings under an explicit coupling-and-instructional-leadership framework (Section 2) that none of the individual studies applies.
Second, relative to prior UAE leadership scoping reviews (; ), which map the breadth of UAE leadership research, this review is narrower in scope (supervision specifically) but deeper in theoretical integration and is, to our knowledge, the first to visually model the UAE's three-authority system in a single figure and to triangulate the UAE evidence base against meta-analytic international benchmarks (Table 7).
Third, relative to international supervision and evaluation scholarship (; ; ; ), this review's contribution is contextual: it tests whether internationally derived claims about coaching-oriented, formative supervision hold in a multi-regulatory, culturally diverse, reform-intensive system, and identifies where UAE evidence is currently too thin (Sections 4.5, 6.5, 7.6) to confirm that transfer.
Fourth, relative to single-country comparative accounts of supervision (; ; ), this review's contribution is to position the UAE explicitly within that comparative landscape (Section 5, Table 6) rather than treating it as a self-contained case — a positioning move that generates the falsifiable claim, absent from the UAE literature reviewed here, that the UAE's accountability-development tension stems from adopting a high-stakes inspection logic without either Finland's trust-based compensating structure or Singapore's competency-track compensating structure.
Collectively, these four contributions reposition supervision in the UAE from a locally described administrative practice into a theoretically grounded, internationally benchmarked case that clarifies not only what UAE supervision currently does, but why its design choices produce the outcomes documented throughout this review.
11 Conclusion
Effective reform will require careful differentiation between summative and formative purposes, investment in evaluator capacity, and governance frameworks that ensure ethical and transparent use of emerging technologies. Framed through loose-coupling and instructional leadership theory, the central policy implication of this review is that UAE supervision reform should not aim for uniform tight coupling across all three layers, but for a deliberate, explicitly stated coupling strategy: tighter at the classroom layer, where feedback and coaching directly shape teaching quality — a linkage supported by the meta-analytic evidence in Table 7 — and more selectively coordinated at the governance layer, where shared standards and data definitions matter more than procedural uniformity.
11.1 Evidentiary Status of this review’s claims
To avoid conflating different strengths of evidence, the claims made throughout this review are separated below into four evidentiary tiers: (a) what UAE evidence directly demonstrates, (b) what the theoretical analysis suggests but does not itself test, (c) what international evidence makes plausible for the UAE without UAE-specific confirmation, and (d) what requires UAE-specific testing before it can be treated as established. Table 11 assigns each major claim in this review to one tier.
Table 11
| Tier | What it means | Claims in this review assigned to this tier |
|---|---|---|
| (a) UAE-demonstrated | Directly observed in UAE data reviewed here (Table 4, §4) | Regulatory fragmentation across MOE/ADEK/KHDA reporting (§4.4); evaluator-capacity gaps and psychological-safety deficits reported by UAE teachers/principals (§7.3, §7.5); concentration of the UAE evidence base in Abu Dhabi public schools and overlapping author teams (§3.4, §4.5) |
| (b) Theory-suggested | Follows from loose-coupling/instructional leadership theory (§2) but not itself directly tested in the UAE | That the three-authority structure produces predictable, tiered fragmentation rather than aberrant dysfunction (§2.1, §4.2); that classroom-level coupling should be tighter than governance-level coupling for supervision to function as an improvement lever (§2.2, §2.3) |
| (c) International-plausible | Supported by international meta-analytic/comparative evidence (Tables 7–9) but not confirmed with UAE data | That coaching-oriented supervision with cognitive-modeling training would improve UAE instructional quality (§6.6, §9.1); that a Singapore-style competency-track layer would convert appraisal into development (§5.3, §9.6); that AI-assisted feedback could improve consistency if the Table 9 safeguards are met (§9.2) |
| (d) Requires UAE-specific testing | Cannot be resolved by theory or international evidence alone; needs a UAE study | Whether tighter MOE-ADEK-KHDA coherence measurably improves outcomes (§9.4); whether digital evidence systems improve instructional outcomes rather than merely digitizing existing ratings (§9.3); whether any AI-assisted supervision pilot meeting the Table 9 safeguards actually improves feedback quality (§9.2); whether a piloted career-track layer improves UAE principal/teacher outcomes (§9.6) |
Evidentiary tier assigned to each major claim made in this review, distinguishing what UAE data directly show from what theory, international evidence, or future UAE testing would be needed to establish.
11.2 Implications for policymakers and school leaders
Four implications follow directly from the evidence synthesized above, ordered from the most to the least directly evidenced.
State the coupling strategy explicitly. MOE, ADEK, and KHDA currently share standards vocabulary without a published statement of how tightly each governance function should be coupled to classroom practice; Sections 4 and 8 (Figure 1) suggest this ambiguity, not the existence of three regulators per se, is the structural driver of the fragmentation documented in Section 7.
Prioritize evaluator cognitive-modeling training over additional documentation infrastructure. The single largest moderator identified in the international coaching literature reviewed here is supervisors making their pedagogical reasoning explicit and demonstrable (d = 0.90; ) — a training investment, not a reporting-system investment.
Commission independent evaluation of ADEK and KHDA policy effectiveness. Section 4.5 and Table 4 show that the evidence base for both authorities currently rests on their own self-published reporting; an externally conducted evaluation, of the kind Table 7's international benchmarks are built on, would materially strengthen the evidence available for future policy decisions.
Pilot, rather than scale, a Singapore-style career-track layer. Section 9.6 treats this as a conditional recommendation precisely because no UAE trial exists; a small-scale pilot with a pre-registered evaluation design would generate the UAE-specific evidence this review's evidence base currently lacks.
11.3 Directions for future research
The critical synthesis in Sections 4.5, 5.6, 6.5, and 7.6 points to a consistent research gap: UAE supervision research is almost entirely qualitative, concentrated in Abu Dhabi public schools, and produced by a small number of overlapping author teams. Three research priorities follow. First, a quantitative, adequately powered study linking supervisory practice to measured instructional or student outcomes in the UAE would close the gap this review repeatedly flags against the international meta-analytic benchmarks in Table 7. Second, research extending beyond Abu Dhabi public schools — into MOE federal schools and Dubai's private sector — is needed before the practice-level claims synthesized here can be treated as system-wide rather than zone-specific. Third, a GCC-wide comparative study, building on the regional workforce-composition evidence in Section 5.5, would test whether the coupling model proposed in this review (Figure 1) generalizes beyond the UAE to the wider Gulf region.
This conclusion should be read with the limitations set out in Section 3.4 in mind: it rests on a narrative Policy and Practice Review of a UAE evidence base that is concentrated among a small number of overlapping author teams, weighted toward regulator-published policy documents over independent evaluations, and, per the evidence-basis statements in Section 9, only partially UAE-tested. The recommendations above are accordingly offered as a comparatively and theoretically grounded policy direction, calibrated against the strongest available international meta-analytic evidence, rather than as a causally established program for the UAE specifically. Their strongest, most directly evidenced elements are the diagnosis of the coupling problem (Sections 4–7) and the confirmation, via meta-analytic triangulation (Table 7), that the mechanisms UAE studies identify are directionally consistent with international evidence at scale. The specific innovations proposed to resolve the coupling problem (Section 9) are more speculative and should be piloted with explicit evaluation designs before scale adoption.
Statements
Author contributions
HR: Conceptualization, Methodology, Formal analysis, Project administration, Supervision, Writing – original draft, Writing – review & editing. TA: Conceptualization, Investigation, Data curation, Writing – original draft, Writing – review & editing. HB: Investigation, Data curation, Validation, Writing – review & editing. KH: Formal analysis, Resources, Writing – review & editing. WM: Investigation, Visualization, Writing – review & editing.
Funding
The author(s) declared that financial support was not received for this work and/or its publication.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was used in the creation of this manuscript. The authors declare that generative AI tools were used solely for language editing, paraphrasing, and formatting purposes, including grammar correction and reference organization. No AI tools were used to generate original research content, data analysis, or interpretations. The authors have reviewed and verified all content and take full responsibility for the accuracy, integrity, and originality of the manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
1
AbdallahA. K.Al KaabiA.QablanA.MalakoutiM. (2025). Instructional leadership in the UAE: a systematic review and thematic synthesis. Educ. Manag. Adm. Leadersh., 1–25. 10.1177/17411432251406021
2
Abu Dhabi Department of Education and Knowledge (2024). ADEK School Quality Assurance Policy (Version 1.1). Abu Dhabi: Abu Dhabi Department of Education and Knowledge. Available online at:https://www.adek.gov.ae(Accessed April 27, 2026).
3
AlanogluM. (2022). The role of instructional leadership in increasing teacher self-efficacy: a meta-analytic review. Asia Pac. Educ. Rev.23 (2), 233–244. 10.1007/s12564-021-09726-5
4
AlghamdiL. H.AlghizziT. M. (2025). Educators’ reflections on AI-automated feedback in higher education: a structured integrative review of potentials, pitfalls, and ethical dimensions. Front. Educ.10, 1704820. 10.3389/feduc.2025.1704820
5
AlharthiM. (2024). Teacher education programs in the Arabian gulf through the eyes of novice teachers: a systematic review with narrative synthesis. Gulf Education and Social Policy Review (GESPR)5 (2) 118–137, 10.18502/gespr.v5i2.16403
6
AlkaabiA. M. (2025). A qualitative multi-case study of supervision in the principal evaluation process in the United Arab Emirates. Int. J. Leadersh. Educ.28. 53–8010.1080/13603124.2021.2000032
7
AlkaabiA. M.AbdallahA. K. (2024). Portfolio practices in the principal evaluation process: a qualitative case study. Heliyon10 (19), e39467. 10.1016/j.heliyon.2024.e39467
8
AlkaabiA. M.AbdallahA. K.AbdallahR.Al HasanM.MohamedA. (2025). Principal engagement in the teacher evaluation process in remote schools: a narrative inquiry of language teachers. Forum Linguist. Stud.7 (3), 209–230. 10.30564/fls.v7i3.8502
9
AlkaabiA. M.AlmaamariS. A. (2020). Supervisory feedback in the principal evaluation process. International Journal of Evaluation and Research in Education (IJERE)9 (3), 503–509. 10.11591/ijere.v9i3.20504
10
Al MaktoumS. B.Al KaabiA. M. (2024). Exploring teachers’ experiences within the teacher evaluation process: a qualitative multi-case study. Cogent Educ.11 (1), 2287931. 10.1080/2331186X.2023.2287931
11
AlmheiriA. S. H.AbuhassnaH. (2024). Exploring the adoption of cutting-edge management practices by school principals. International Journal of Evaluation and Research in Education (IJERE)13 (4), 2095–2107. 10.11591/ijere.v13i4.28759
12
AlShehhiF.AlzouebiK. (2020). The hiring process of principals in public schools in the United Arab Emirates: practices and policies. Int. J. Educ. Literacy Stud.8 (1), 74–83. 10.7575/aiac.ijels.v.8n.1p.74
13
DemszkyD.LiuJ.HillH. C.JurafskyD.PiechC. (2024). Can automated feedback improve teachers’ uptake of student ideas? Evidence from a randomized controlled trial in a large-scale online course. Educ. Eval. Policy. Anal.46, 483–50510.3102/01623737231169270
14
EhrenM. C. M.GustafssonJ.-E.AltrichterH.SkedsmoG.KemethoferD.HuberS. G. (2015). Comparing effects and side effects of different school inspection systems across Europe. Comp. Educ.51 (3), 375–400. 10.1080/03050068.2015.1045769
15
FusarelliL. D. (2002). Tightly coupled policy in loosely coupled systems: institutional capacity and organizational change. J. Educ. Adm.40 (6), 561–575. 10.1108/09578230210446045
16
GlickmanC. D.GordonS. P.Ross-GordonJ. M. (2018). SuperVision and Instructional Leadership: A Developmental Approach. 10th ed. Boston, MA: Pearson.
17
GrissomJAEgaliteAJLindsayCA (2021). How Principals Affect Students and Schools: A Systematic Synthesis of two Decades of Research. The Wallace Foundation. Available online at:http://www.wallacefoundation.org/principalsynthesis(Accessed April 27, 2026).
18
HallingerP. (2011). Leadership for learning: lessons from 40 years of empirical research. J. Educ. Adm.49 (2), 125–142. 10.1108/09578231111116699
19
HallingerP.HeckR. H. (1996). Reassessing the principal’s role in school effectiveness: a review of empirical research, 1980–1995. Educ. Adm. Q.32 (1), 5–44. 10.1177/0013161X96032001002
20
HojeijZ.AlsuwaidiS.AhmedS. (2024). Scoping the literature on professional development for educators and educational leaders in the UAE. Res. Educ. Adm. Leadersh.9 (3), 216–251. 10.30828/real.1480120
21
HojeijZ.IbrahimA.BaroudiS. (2023). Stressors and stress coping strategies of UAE school leaders. Int. J. Manag. Educ.17 (4), 339–360. 10.1504/IJMIE.2023.131218
22
JuhjiJ.Ma'murI.NugrahaE.NurhadiA.TarihoranN. (2022). A meta-analysis study of principal leadership and teacher job satisfaction. Al-Ishlah J. Pendidik.14 (2), 1645–1652. 10.35445/alishlah.v14i2.1498
23
KaneT. J.TaylorE. S.TylerJ. H.WootenA. L. (2011). Evaluating teacher effectiveness: can classroom observations identify practices that raise achievement?Education Next11 (3), 54–61.
24
Knowledge and Human Development Authority (2017). United Arab Emirates School Inspection Framework. Dubai: Knowledge and Human Development Authority. Available online at:https://web.khda.gov.ae/en/Resources/Publications/School-Inspection/UAE-School-Inspection-Framework-2015-16(Accessed April 27, 2026).
25
Knowledge and Human Development Authority (2024). Inspection key Findings 2023-2024. Available online at:https://web.khda.gov.ae/en/Resources/Publications/School-Inspection/Inspection-Key-Findings-2023-2024(Accessed April 27, 2026).
26
KraftM. A.BlazarD.HoganD. (2018). The effect of teacher coaching on instruction and achievement: a meta-analysis of the causal evidence. Rev. Educ. Res.88 (4), 547–588. 10.3102/0034654318759268
27
KraftM. A.GilmourA. F. (2016). Revisiting the widget effect: teacher evaluation reforms and the distribution of teacher effectiveness. Educ. Res.45 (4), 234–249. 10.3102/0013189X16642222
28
LeithwoodK.HarrisA.HopkinsD. (2020). Seven strong claims about successful school leadership revisited. Sch. Lead. Manag.40 (1), 5–22. 10.1080/13632434.2019.1596077
29
LiebowitzD. D. (2022). Teacher evaluation for growth and accountability: under what conditions does it improve student outcomes?Harv. Educ. Rev.92 (4), 533–565. 10.17763/1943-5045-92.4.533
30
LiebowitzD. D.PorterL. (2019). The effect of principal behaviors on student, teacher, and school outcomes: a systematic review and meta-analysis of the empirical literature. Rev. Educ. Res.89 (5), 785–827. 10.3102/0034654319866133
31
LitzD. R. (2014) Perceptions of school leadership in the United Arab Emirates (UAE). Doctoral thesis. University of Calgary. PRISM. 10.11575/PRISM/27288
32
MacabuhayH. (2026). Clinical supervision on the professional development of teachers. Int. J. Sustain. Adv. Integr. Res.2 (2) 831–837. 10.65339/ijsair.V2.I2.260
33
MetteI. M.RangeB. G.AndersonJ.HvidstonD. J.NieuwenhuizenL.DotyJ. (2017). The wicked problem of the intersection between supervision and evaluation. Int. Electron. J. Elem. Educ.9 (3), 709–724.
34
Ministry of Education (2017). UAE School Inspection Framework. Abu Dhabi: Ministry of Education. Available online at:https://www.moe.gov.ae(Accessed April 27, 2026).
35
Ministry of Education (n.d.). Abu Dhabi: Ministry of Education. Teacher Licensing System. Available online at:https://tls.moe.gov.ae(Accessed April 27, 2026).
36
MokS. Y.StaubF. C. (2021). Does coaching, mentoring, and supervision matter for pre-service teachers’ planning skills and clarity of instruction? A meta-analysis of (quasi-)experimental studies. Teach. Teach. Educ.107, 103484. 10.1016/j.tate.2021.103484
37
NgP. T.WongB. (2019). “Introduction: school leadership and educational change in Singapore,” in School Leadership and Educational Change in Singapore, eds. WongB.HaironS.NgP. T. (Cham: Springer), 1–6. 10.1007/978-3-319-74746-0_1
38
NugentG.HoustonJ.KunzG.ChenD. (2023). Analysis of instructional coaching: what, why and how. Int. J. Mentor. Coach. Educ.12, 402–423. 10.1108/IJMCE-08-2022-0066
39
OECD (2016). School Leadership for Learning: Insights from TALIS 2013. Paris: OECD Publishing. 10.1787/9789264258341-en
40
OrtonJ. D.WeickK. E. (1990). Loosely coupled systems: a reconceptualization. Acad. Manag. Rev.15 (2), 203–223. 10.5465/amr.1990.4308154
41
PainoM. (2018). From policies to principals: tiered influences on school-level coupling. Soc. Forces96 (3), 1119–1154. 10.1093/sf/sox075
42
PerrymanJ.BradburyA.CalvertG.KilianK. (2025). ‘A tipping point’ in teacher retention and accountability: the case of inspection. Br. J. Educ. Stud.73 (2), 181–200. 10.1080/00071005.2024.2439791
43
QuA.WenY.ZhangJ.WenY.ZhaoY.PrakashA.et al (2025). Classmind: scaling classroom observation and instructional feedback with multimodal AI. arXiv:2509.18020[Preprint]. 10.48550/arXiv.2509.18020
44
RaiJ.Beresford-DeyM. (2025). School leadership in the United Arab Emirates: a scoping review. Educ. Manag. Adm. Leadersh.53 (6), 1335–1357. 10.1177/17411432231218129
45
RidgeN.ShamiS.KippelsS.FarahS. (2014). Expatriate Teachers and Education Quality in the Gulf Cooperation Council. Ras Al Khaimah: Sheikh Saud bin Saqr Al Qasimi Foundation for Policy Research.
46
RobinsonV. M. J.LloydC. A.RoweK. J. (2008). The impact of leadership on student outcomes: an analysis of the differential effects of leadership types. Educ. Adm. Q.44 (5), 635–674. 10.1177/0013161X08321509
47
SahlbergP. (2015). Finnish Lessons 2.0: What can the World Learn from Educational Change in Finland?2nd ed. New York, NY: Teachers College Press.
48
ShenJ.WuH. (2025). The relationship between principal leadership and student achievement: a multivariate meta-analysis with an emphasis on conceptual models and methodological approaches. Educ. Adm. Q.61 (2), 234–281. 10.1177/0013161X241286527
49
SteinbergM. P.DonaldsonM. L. (2016). The new educational accountability: understanding the landscape of teacher evaluation reform. Educ. Res.45 (2), 67–74. 10.3102/0013189X16630898
50
SteinbergM. P.SartainL. (2015). Does better observation make better teachers? New evidence from a teacher evaluation pilot in Chicago. Education Next15 (1), 70–76.
51
SunJ.ZhangR.MurphyJ.ZhangS. (2024). The effects of academic press on student learning and its malleability to school leadership: a meta-analysis of 30 years of research. Educ. Adm. Q.60 (2), 226–268. 10.1177/0013161X231217226
52
UNESCO (2024). Global Education Monitoring Report 2024/5: Leadership in Education — Lead for Learning. Paris: UNESCO.
53
WeickK. E. (1976). Educational organizations as loosely coupled systems. Adm. Sci. Q.21 (1), 1–19. 10.2307/2391875
54
WuH.ShenJ. (2022). The association between principal leadership and student achievement: a multivariate meta-meta-analysis. Educ. Res. Rev.35, 100423. 10.1016/j.edurev.2021.100423
Summary
Keywords
educational supervision, instructional leadership, loose-coupling theory, policy reform, quality assurance, school leadership, teacher evaluation, United Arab Emirates
Citation
Badawy HRI, Alshloul T, Badawy HRI, Hassan Almaazmi KM and Mohsen W (2026) Supervision in educational leadership: practices, challenges, and innovations in the UAE — a multi-level coupling analysis of the current and future. Front. Educ. 11:1862260. doi: 10.3389/feduc.2026.1862260
Received
22 April 2026
Revised
29 July 2026
Accepted
31 July 2026
Published
14 August 2026
Volume
11 - 2026
Edited by
Mohammed Borhandden Musah, Emirates College for Advanced Education, United Arab Emirates
Reviewed by
Rany Sam, National University of Battambang, Cambodia
Mark-Jhon Prestoza, Isabela State University - Cauayan Campus, Philippines
Updates
Copyright
© 2026 Badawy, Alshloul, Badawy, Hassan Almaazmi and Mohsen.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Hosam R. I. Badawy 201970119@uaeu.ac.ae
† These authors have contributed equally to this work
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.