Abstract
Background:
The successive releases of GPT-3 (May 2020) and ChatGPT (November 2022) have been widely hypothesized to constitute inflection points in the automation of cognitive labor. Yet empirical evidence distinguishing AI-driven displacement from secular trends, pandemic disruption, and cyclical variation has remained fragmented and geographically narrow.
Methods:
Following PRISMA 2020 guidelines, we systematically searched six academic databases (Scopus, Web of Science, EconLit, SSRN, IEEE Xplore, Google Scholar) for empirical studies documenting observed—not predicted—labor market changes since 2020. From 1,847 initial records, 94 studies meeting inclusion criteria were retained for qualitative synthesis and 42 for quantitative data extraction.
Results:
Across synthesized studies, converging evidence documents: (1) a 14–41% reduction in postings for entry- and mid-level software development and content-creation roles in high-income economies between 2022 and 2024 (range across individual studies: −14% to −41%; median: −23%); these figures are not pooled estimates but represent the span observed across non-overlapping study designs and geographies, and should be interpreted as illustrative of the order of magnitude of the effect rather than as a meta-analytic point estimate. (2) a 15%–22% wage premium for workers demonstrating AI-augmentation capabilities; (3) heterogeneous sectoral effects, with infrastructure, security, and quality-assurance roles expanding alongside developer role contraction; and (4) evidence from online labor markets of a 2%–21% reduction in posting volumes for automatable creative tasks following ChatGPT's release. Wage polarization, credential erosion, and geographic unevenness characterize the aggregate pattern.
Conclusions:
Observable labor market data, while constrained by short observation windows, already document patterns consistent with AI-driven displacement rather than mere transformation—concentrated among routine cognitive tasks and junior roles, with preliminary but material evidence that developing economies reliant on cognitive services outsourcing face disproportionate disruption through both direct exposure and indirect demand-erosion channels. The displacement is concentrated among routine cognitive tasks and junior roles, with developing economies potentially facing disproportionate disruption. Persistent data gaps—especially concerning worker-level outcomes, informal labor, and non-Anglophone markets—warrant urgent research investment.
1 Introduction
Debates about technology and employment have accompanied every major wave of industrial transformation, from the mechanization of textile production in the early nineteenth century to the computerization of clerical work in the late twentieth. In each case, the theoretical prediction of permanent, large-scale unemployment proved premature: labor markets adapted, new task demands emerged, and employment levels recovered over medium horizons (; ). This historical pattern has generated widespread skepticism about contemporary claims that artificial intelligence represents a qualitatively different threat to employment ().
The release of OpenAI's GPT-3 in May 2020 and, more consequentially, ChatGPT in November 2022, shifted the locus of the debate from prediction to observation. For the first time, AI systems capable of generating code, prose, analysis, and creative content at human-competitive quality were placed in the hands of hundreds of millions of users and embedded directly into enterprise workflows. The question was no longer whether AI could automate cognitive tasks in principle, but whether observable labor market outcomes had begun to change in measurable ways.
As of early 2026, a body of empirical evidence has accumulated across job posting databases, platform labor market studies, firm-level surveys, and national labor force data. This evidence has not been systematically synthesized. Existing reviews either focus on predictive frameworks (; ) or examine pre-LLM automation (; ). The present review fills this gap by systematically assembling and evaluating the empirical record of observed, post-GPT-3 labor market changes attributable to or consistent with AI-driven displacement.
We emphasize the distinction between prediction and observation. Predictive studies—which estimate the share of tasks or occupations susceptible to automation based on task content analysis (; )—are not the primary object of this review. We instead focus on studies that measure what has already happened: changes in job posting volumes, wages, employment levels, task composition, and worker outcomes that postdate LLM deployment. This evidential standard is more demanding, given shorter observation windows and stronger identification challenges, but it provides the empirical foundation that the field requires.
1.1 Scope and research questions
This review addresses four research questions derived from the theoretical literature on AI and labor:
RQ1: What empirical evidence documents changes in the volume and composition of job postings, employment, or task demand attributable to LLM deployment since 2020?
RQ2: What evidence documents changes in wages, compensation, or skill premiums associated with AI adoption?
RQ3: How do observed displacement effects vary across sectors, occupations, skill levels, and geographies?
RQ4: What are the principal methodological limitations of existing empirical work, and what gaps warrant future research?
Several key terms used throughout this review require explicit disambiguation, as they are used inconsistently in the wider literature.
Labor displacementrefers here to the reduction in demand for specific categories of human labor, measurable through job posting declines, employment level changes, or platform task volume reductions, and does not presuppose permanent or aggregate unemployment.
Technological unemploymentrefers to the broader hypothesis that AI-driven productivity gains exceed the pace of new task and role creation, producing net employment reduction at the economy-wide level—a claim that goes beyond what the current evidence supports.
Wage polarizationrefers to the simultaneous growth in compensation at the top of the wage distribution and compression or decline at the middle, consistent with task-based theoretical predictions (
) and documented in Section
4.2.3; it does not imply uniform wage decline. These distinctions matter for evidence interpretation: the studies synthesized here most directly support the first construct; they provide suggestive but incomplete evidence for the second; and they document the third as an emerging pattern consistent with, but not yet causally established by, LLM deployment.
2 Methods
This review follows the PRISMA 2020 reporting guidelines () as the most widely adopted systematic review reporting standard, applied here to an economic and labor market context rather than a clinical one. This section describes search strategy, eligibility criteria, data extraction procedures, and synthesis approach.
2.1 Eligibility criteria
Studies were eligible for inclusion if they satisfied all of the following criteria:
Empirical scope: The study presented original empirical findings (not purely theoretical or predictive) on labor market outcomes including job postings, employment levels, wages, task composition, platform work volumes, or related measures.
Temporal scope: Data covered some portion of the period from January 2020 to January 2026, capturing at least the GPT-3 or ChatGPT release as a potential treatment event.
Causal relevance: The study either (a) used quasi-experimental or interrupted time series methods to isolate AI effects, (b) measured AI adoption as a covariate alongside outcome measures, or (c) documented systematic trends in AI-exposed occupations or platforms with explicit comparators.
Language: Published in English.
Peer review status: Peer-reviewed journal articles, working papers posted on recognized repositories (NBER, SSRN, IZA, arXiv), or credible institutional reports (McKinsey Global Institute, OECD, World Bank) meeting pre-specified quality criteria.
Studies were excluded if they: (a) relied exclusively on expert surveys or simulated automation propensity scores without labor market outcome data; (b) examined robotic process automation or pre-LLM AI without distinguishing these from LLM effects; (c) covered only pre-2020 periods; or (d) reported only country-level aggregate unemployment statistics without occupational or sectoral decomposition.
2.2 Search strategy
We searched six databases on January 15, 2026: Scopus, Web of Science, EconLit, SSRN, IEEE Xplore, and Google Scholar. Search terms combined three conceptual clusters using Boolean operators:
Cluster A (AI technology): “large language model” OR “LLM” OR “ChatGPT” OR “GPT-3” OR “GPT-4” OR “generative AI” OR “AI coding assistant” OR “GitHub Copilot”
Cluster B (labor outcomes): “job displacement” OR “employment” OR “job posting” OR “wage” OR “labor market” OR “technological unemployment” OR “freelance” OR “gig work” OR “task demand”
Cluster C (empirical qualifier): “empirical” OR “observed” OR “evidence” OR “natural experiment” OR “interrupted time series” OR “difference-in-differences”
We additionally searched the reference lists of all included studies (“snowballing”) and manually checked the publications lists of key research groups at MIT, Stanford, NBER, Oxford, and the IZA Institute. Grey literature from McKinsey Global Institute, World Economic Forum, LinkedIn Economic Graph, Burning Glass/Lightcast, and the International Labour Organization was searched separately.
2.3 Screening and selection
Records retrieved from database searches were de-duplicated using Covidence systematic review software. Two independent reviewers screened titles and abstracts for eligibility, with disagreements resolved through discussion and, where necessary, adjudication by a third reviewer. Inter-rater agreement at title/abstract screening reached κ = 0.81, indicating strong concordance. Full texts of potentially eligible records were retrieved and assessed against eligibility criteria. Table 1 summarizes the PRISMA flow. The complete PRISMA 2020 screening and selection process is presented in Figure 1.
Table 1
| Stage | Records (n) |
|---|---|
| Database search results (combined) | 1,847 |
| Records after de-duplication | 1,412 |
| Records screened (title/abstract) | 1,412 |
| Records excluded at title/abstract | 1,089 |
| Full texts assessed for eligibility | 323 |
| Full texts excluded (with reasons) | 229 |
| Not empirical/predictive only | 88 |
| Pre-2020 data only | 51 |
| No AI-exposure variable or comparator | 47 |
| Aggregate unemployment data only | 28 |
| Language/access issues | 15 |
| Studies included in qualitative synthesis | 94 |
| Studies included in quantitative extraction | 42 |
PRISMA 2020 flow diagram summary.
Figure 1
2.4 Data extraction and quality assessment
For each included study, two reviewers independently extracted: (a) study design and identification strategy, (b) geographic and sectoral scope, (c) temporal window and AI exposure operationalization, (d) primary outcome measures and effect sizes, (e) controls and confounders addressed, and (f) key limitations. Effect sizes were harmonized to percentage changes from baseline where possible. Where studies reported only regression coefficients, marginal effects were calculated at sample means.
Quality was assessed using a modified Newcastle-Ottawa Scale adapted for observational economic studies and an additional five-item checklist specific to AI-exposure studies: (1) clear definition of AI treatment event or adoption measure; (2) pre-treatment baseline documented; (3) comparator group specified; (4) confounders (COVID-19, economic cycle, platform-specific shocks) addressed; (5) sensitivity or robustness checks reported. Studies scoring ≥3 of 5 criteria were classified as high quality; 2 of 5 as moderate; ≤1 as low. Low-quality studies were included in narrative synthesis but excluded from quantitative pooling.
2.5 Synthesis approach
Given substantial heterogeneity in study designs, outcome measures, and geographic contexts, a quantitative meta-analysis with pooled effect sizes was not feasible. We employed narrative synthesis structured around the four research questions, supplemented by tabular presentation of quantitative estimates and vote-counting for directional consensus across studies. We report ranges and medians of effect sizes rather than pooled estimates, and we use traffic-light summary tables to signal evidence strength for each primary finding.
3 Theoretical context and conceptual framework
Before presenting findings, we briefly situate the review within the theoretical frameworks most directly implicated by LLM deployment.
3.1 What makes LLMs theoretically distinct
Standard task-based models of automation (
A scope clarification is warranted. This review focuses specifically on large language models and generative AI systems, defined as AI architectures capable of generating natural language, code, or multimodal content at human-competitive quality, with GPT-3 (May 2020) and ChatGPT (November 2022) as primary periodization markers. We do not systematically review evidence on robotic process automation, computer vision systems, recommendation algorithms, or earlier generation machine learning tools, except where included studies use these as comparison conditions or pre-treatment baselines. This boundary is material: prior automation literature (
The Turing Trap hypothesis (
3.2 Channels of displacement, restructuring, and validation
The empirical literature, as we shall see, suggests that LLM-driven displacement operates through at least four channels: (1) direct task replacement, where AI systems perform tasks previously assigned to workers; (2) productivity augmentation enabling workforce reduction, where AI raises individual productivity enough that fewer workers are needed for the same output; (3) role consolidation, where AI enables senior workers to absorb functions previously distributed across junior and mid-level staff; and (4) demand erosion, where the availability of AI tools reduces client demand for human professionals in relevant services.
A fifth channel, which the empirical record in Section 4.3.5 elevates to theoretical standing, is validation and quality assurance restructuring: as AI systems take on creation tasks (code generation, content drafting, data synthesis), the nature of human oversight shifts from routine checking against specifications to non-routine cognitive evaluation of AI-generated outputs for correctness, safety, and hallucination—a qualitatively distinct task bundle that generates new labor demand even as it emerges from displacement in adjacent creation roles. This channel is conceptually captured by the paper's title: AI drives a labor market reorganization across a three-stage cycle of creation (now increasingly AI-performed), validation (increasingly human-critical), and obsolescence (the career pathways and curricula premised on the prior division of labor). These channels are not mutually exclusive and are often empirically difficult to disentangle.
4 Results
We organize findings according to the four research questions, presenting evidence from job posting analyses, online platform labor markets, firm-level studies, and macroeconomic data. Following peer review, one additional study (
Table 2
| Study (year) | Geography | Method | Period | Primary outcome | Effect size | Quality |
|---|---|---|---|---|---|---|
| Online (global) | DiD, Upwork platform | 2021–2023 | Freelance writing/coding postings | −21% (writing) post-ChatGPT | High | |
| OECD | Cross-sectional DiD | 2022–2024 | LLM-exposed occupation wages | +4.8% wage premium (AI-skilled workers) | High | |
| USA | BLS + firm survey | 2018–2024 | Employment share, AI-exposed sectors | −3.5% employment, AI-intensive firms | Moderate | |
| USA | O*NET + job posting panel | 2019–2024 | Posting volumes by AI exposure | −15% to −34% (high-exposure roles) | High | |
| USA | RCT (GitHub Copilot) | 2022 | Developer productivity | +55.8% task completion speed | High | |
| USA | ADP payroll panel (millions of workers) | 2021–2025 | Early-career employment, AI-exposed occupations | −13% (ages 22–25, highest AI exposure quintile) | High | |
| USA (MTurk) | RCT | 2022–2023 | Writing task output quality + time | +37% quality; −40% time | High | |
| USA (BCG) | RCT | 2023 | Consultant task performance | +12.2% quality score (AI users) | High | |
| USA | Lightcast job posting panel | 2020–2024 | Posting volumes: content creation | −23% (copywriting/editing) | Moderate | |
| USA, EU | GitHub + job boards | 2022–2024 | Junior dev postings | −31% to −38% (entry-level coding) | Moderate | |
| Online (AMT) | Platform data | 2021–2024 | Data annotation task volumes | −18% (automatable annotation) | Moderate | |
| OECD 18 countries | DiD, LinkedIn data | 2022–2024 | AI-adjacent vs. AI-exposed roles | +26% (AI-adjacent), −14% (AI-replaced) | High | |
| Poland | Job posting panel | 2018–2024 | IT sector postings | −41% (2023–2024 vs. 2021–2022) | Moderate | |
| Global (121 countries) | Labour Force Surveys | 2019–2023 | Clerical occupation employment | −3.4% (high-income), +0.2% (low-income) | High | |
| Global | Occupational survey | 2022–2023 | Roles hiring for AI augmentation | +28% (demand for AI-skilled workers) | Moderate | |
| EU-27 | Job posting analysis | 2021–2024 | AI skill mentions in postings | +340% (AI tool mentions in requirements) | Moderate |
Summary of primary quantitative findings from included studies (selected).
DiD, difference-in-differences; ITS, interrupted time series; RCT, randomized controlled trial; BLS, Bureau of Labor Statistics.
Figure 2 presents these effect sizes visually, ordered by direction and magnitude, with quality ratings indicated by color. Ranges are shown where studies reported low–high bounds rather than point estimates.
Figure 2

Sectoral heterogeneity in AI-driven labor market effects. Blue = declining; gray = stable; green = growing.
4.1 RQ1: evidence on job posting volumes, employment, and task demand
4.1.1 Job posting analyses
The most direct and abundant evidence comes from longitudinal analyses of job posting databases, which offer near-real-time measurement of employer demand with fine occupational and temporal resolution.
For the United States,
4.1.2 Online platform labor markets
Online freelance platforms offer a complementary empirical window because their volume and wage data are more granular and more rapidly responsive to AI adoption than economy-wide employment data
Taken together, online platform evidence suggests that AI-driven displacement is observable at the margin of the labor market—affecting lower-skill, lower-wage, more routine cognitive work—before it penetrates the core employment relationship. This sequencing has precedent in prior automation waves (
4.1.3 Macroeconomic and administrative data
Aggregate labor force survey data from the 2020–2025 period present a more ambiguous picture, partly because LLM deployment is still recent and its labor market effects may not yet be captured in employment stocks (as opposed to flows).
The
4.2 RQ2: wage effects and skill premiums
4.2.1 Aggregate salary trends in AI-affected markets
A recurring and initially counterintuitive finding across multiple studies is that average advertised salaries have increased in markets experiencing the largest job posting declines. This pattern is consistent with compositional shift: as junior and mid-level positions are eliminated, the remaining postings skew toward senior, specialized roles commanding higher compensation.
The salary-to-experience ratio remains stable across eras, suggesting that per-unit-of-experience compensation has not changed dramatically; rather, the market selects for workers with substantially more experience.
4.2.2 AI skill premiums
Several studies document a positive and growing wage premium for workers demonstrating AI-tool proficiency.
4.2.3 Wage polarization
The combined evidence is consistent with growing wage polarization: a simultaneous increase in wages at the top of the distribution (driven by premium compensation for AI-augmented, senior roles) and compression at the middle and lower end of AI-exposed skill categories. This pattern aligns with
4.3 RQ3: heterogeneity of effects
4.3.1 Sectoral variation
The aggregate contraction in AI-exposed roles masks substantial sectoral heterogeneity. Table 3 summarizes the sectoral findings across included studies.
Table 3
| Sector | Direction | Magnitude (range) | Primary mechanism | Key studies |
|---|---|---|---|---|
| Software development (general) | ↓ Declining | 14%–41% posting decline | Direct task substitution; productivity augmentation | |
| Content creation/copywriting | ↓ Declining | 21%–23% posting decline | Direct text generation substitution | |
| Data annotation/labeling | ↓ Declining | 18% task volume decline | Automatable annotation tasks replaced | |
| Cybersecurity/security engineering | ↑ Growing | +36% (OECD) | Expanded attack surface from AI; AI code vulnerabilities | |
| Management consulting (junior) | ↓ Declining | −8% to −15% (US) | AI-assisted analysis reduces junior analyst headcount | |
| Healthcare IT/clinical | → Stable | −2% to +4% | Regulatory constraints; high non-routine content |
Sectoral heterogeneity in AI-driven labor market effects. .
Magnitudes are not directly comparable across studies due to differences in measurement period, geography, and baseline definition.
The emerging pattern across published studies is a structural shift from software creation toward infrastructure orchestration: roles whose core function is the manual production of code or content are declining, while roles focused on deploying, securing, validating, and managing AI-generated outputs are growing.
Figure 3 summarizes these sectoral heterogeneity findings, illustrating the magnitude ranges of posting changes across the six sectors with sufficient quantitative data for comparison.
Figure 3

Primary quantitative findings from included studies. Effect sizes reflect heterogeneous methods, outcome measures, geographies, and baseline periods and cannot be pooled or directly compared. Cedefop (2024) actual value = + 340%; truncated for readability.
4.3.2 Occupational level and seniority
Across studies, the displacement effect is consistently more pronounced among junior and mid-level workers than senior practitioners.
This seniority gradient is theoretically interpretable through the lens of role consolidation: senior workers, augmented by AI tools, can absorb functions previously distributed across junior and mid-level staff. The implication is that the pipeline of junior talent development may be disrupted, potentially creating future shortages of experienced practitioners—a concern raised by
4.3.3 Geographic variation
High-income economies with mature IT sectors (USA, UK, Germany, Australia) show more moderate aggregate employment effects—partly because AI adoption in these markets generates demand for AI-adjacent roles that partially offsets displaced positions—while middle-income economies with IT employment concentrated in routine cognitive outsourcing are theoretically more vulnerable.
The
That gap, however, does not mean that developing-economy disruption is hypothetical. Substantial evidence now documents concrete, realized displacement in three distinct channels. The first is the business process outsourcing (BPO) sector. India's major IT services firms—TCS, Infosys, and Wipro—collectively reduced their workforces by more than 80,000 positions over 2023–2025 as global clients adopted AI tools that reduced demand for outsourced routine coding, testing, and customer service work (
The second channel is the platform-mediated data annotation and gig labor market. In March 2024, Scale AI's subsidiary Remotasks abruptly terminated operations in Kenya, Nigeria, and Pakistan—countries where the platform had become a primary income source for thousands of young workers engaged in data labeling, image annotation, and content moderation tasks. Workers received termination emails hours before losing access; in Kenya, wages held in third-party payment accounts were frozen and in many cases not recovered (
The third channel is the IMF's cross-country empirical evidence.
Taken together, this evidence base is sufficient to characterize developing-economy disruption not as a hypothesis but as an ongoing, documented phenomenon operating through at least three distinct mechanisms. The assertion that developing economies face disproportionate disruption is supported by substantially more evidence than the single
4.3.4 Gender and demographic dimensions
The
Young workers (aged 18–34) appear disproportionately affected relative to their labor market shares, partly because they disproportionately occupy the entry-level and junior positions most subject to consolidation, and partly because they compete for the lowest-barrier-to-entry online platform tasks most susceptible to LLM substitution.
4.3.5 The QA and security “validation paradox”
A recurring finding across multiple studies—and one that challenges simple displacement narratives—is the simultaneous growth of quality assurance, testing, and security roles alongside declining development and content-creation roles.
This “validation paradox” suggests that AI does not simply replace human labor but restructures it: as AI takes on creation tasks, human labor shifts toward verification, security assessment, and quality control. The QA role in an AI-augmented environment is no longer routine checking against specifications but non-routine cognitive evaluation of AI-generated outputs for correctness, safety, and hallucination—a substantively different task bundle despite the legacy occupational classification.
4.4 RQ4: methodological limitations and research gaps
4.4.1 Identification challenges
The central methodological challenge is identifying AI-specific effects against the confounding backdrop of the COVID-19 pandemic (2020–2022), the macroeconomic cycle, and secular trends in digitalization. The period of LLM deployment overlaps almost entirely with a period of extraordinary labor market volatility, making clean causal identification difficult.
The strongest studies (e.g.,
4.4.2 Measurement scope
Job posting data capture employer demand but not employment levels, hours worked, or worker welfare. A decline in postings could reflect productivity-driven workforce reduction (fewer workers producing the same output), genuine employment contraction, or simply a shift in recruitment channels (from public job boards to AI-assisted targeted outreach).
Salary data in job posting studies are typically sparse, as the majority of postings in most markets do not disclose compensation. Survey-based wage data are less affected by this limitation but may understate short-term adjustment due to survey timing and lag between market changes and data collection.
4.4.3 Geographic and sectoral scope
The empirical literature remains heavily concentrated in high-income, English-speaking economies (particularly the United States). Beyond the
4.4.4 Temporal limitations
The most recent AI milestones—the Agent Era inaugurated by systems like Claude Code (January 2025), Devin, and AutoGPT frameworks capable of autonomous multi-step task completion—fall outside the study period of most included research. If the Agent Era represents a qualitatively new displacement regime characterized by autonomy rather than assistance, the existing literature may systematically understate the medium-term trajectory of displacement.
The short observation window also limits the ability to distinguish between adjustment dynamics and structural change. Labor economists have long documented that large technological shocks produce non-linear employment responses: initial displacement is followed by reabsorption into new tasks and roles, often with a lag of five to ten years (
5 Discussion
5.1 Convergence and divergence across studies
Despite heterogeneity in methods, geographies, and sectors, several findings exhibit strong directional consensus across included studies. The following table presents a traffic-light summary of evidence strength for key claims.
A note on causal inference is warranted before interpreting the patterns below. The included studies vary substantially in their ability to support causal claims. Difference-in-differences designs with pre-treatment parallel trend testing—such as
Table 4 summarizes the convergence and divergence across studies. Throughout this section we use the following convention: findings rated “strong” (✔✔✔) are supported by multiple high-quality causal designs; “moderate” (✔✔) by mixed-quality designs or cross-context replication; “weak” (✔) by single studies or descriptive evidence. Readers should reserve causal language (“AI drove” “AI caused”) for findings in the strong category; for moderate and weak findings, directional language (“consistent with” “associated with”) is the more defensible characterization.
Table 4
| Finding | Direction | Evidence strength | No. studies |
|---|---|---|---|
| Posting declines in LLM-exposed software/content roles (2022–2024) | Negative | Strong (✔✔✔) | 11 |
| Entry-level/junior roles disproportionately affected | Negative | Strong (✔✔✔) | 8 |
| Wage premium for AI-skilled workers | Positive | Moderate (✔✔) | 6 |
| Average advertised salaries increasing despite volume decline | Positive | Moderate (✔✔) | 4 |
| Online platform task volumes declining (writing, annotation) | Negative | Strong (✔✔✔) | 5 |
| Security and validation roles growing | Positive | Moderate (✔✔) | 5 |
| Developing economies experiencing greater disruption | Negative | Moderate (✔✔) | 3 |
| Aggregate employment levels declining in exposed sectors | Negative | Weak (✔) | 4 |
| AI-native specialist roles (data scientist, ML engineer) declining | Negative | Moderate (✔✔) | 3 |
| Infrastructure/DevOps skills growing | Positive | Moderate (✔✔) | 3 |
| Formal degree requirements declining | Negative | Weak (✔) | 2 |
| Onsite work requirements increasing | Mixed | Weak (✔) | 2 |
Evidence traffic-light summary.
✔✔✔, strong convergent evidence from multiple high-quality studies; ✔✔, moderate evidence from mixed-quality studies; ✔, limited evidence or single study.
5.2 Mechanisms of displacement
The four displacement channels identified in Section 3.2—direct task replacement, productivity augmentation, role consolidation, and demand erosion—each receive empirical support, though with different levels of evidence.
Direct task replacement is most clearly documented in online platform markets, where specific task categories (image captioning, text summarization, basic copywriting) have experienced sharp volume declines closely timed with LLM deployment. The
Productivity augmentation reducing headcount receives support from the RCT literature.
Role consolidation is the mechanism most strongly supported by the seniority gradient in displacement effects: the concentration of posting declines among junior and mid-level roles, while senior positions remain relatively stable or grow, is consistent with senior workers absorbing functions previously distributed across larger teams. This mechanism also explains the apparent paradox of rising average salaries alongside falling total posting volumes.
Demand erosion—where AI tools available to clients reduce demand for human professional services—is most clearly implicated in the freelance platform context.
5.3 Comparison with prior automation waves
Several features of the observed LLM displacement pattern differ from prior automation waves in ways that warrant theoretical attention.
First, the speed of the market response appears unprecedented relative to prior automation waves.
Second, the displacement is occurring in cognitive rather than manual domains, and at the high-skill end of the cognitive distribution. Prior automation displaced manufacturing and clerical workers; LLM displacement is most acute among software developers, content creators, and data annotators—a historically protected labor market segment. This inversion of the
Third, the geographic pattern is potentially more regressive than prior automation waves. Industrial robot adoption displaced manufacturing workers primarily in advanced economies (where manufacturing was itself concentrated), with some displacement of offshored manufacturing in developing economies. LLM-driven displacement threatens the cognitive services outsourcing sector—call centers, data annotation, software development, content creation—that represents a primary development pathway for middle-income economies.
5.4 The paradox of AI-native role disappearance
Perhaps the most counterintuitive finding in the emerging empirical literature concerns AI-native roles (data scientist, ML engineer, AI engineer).
5.5 The creation-validation-obsolescence framework
The paper's title frames the observed labor market reorganization as a three-stage cycle that the empirical evidence, taken together, substantiates. AI systems are increasingly performing creation tasks—code generation, content drafting, data annotation—that constituted the core functions of the junior and mid-level roles experiencing the sharpest posting declines documented in Section 4.1. Human labor is not eliminated but restructured toward validation: the quality assurance, security audit, and output verification roles that Section 4.3.5 documents are simultaneously growing. Underlying both dynamics is obsolescence—not of human workers per se, but of the career pathways, educational curricula, and hiring pipelines premised on the prior division of cognitive labor. The entry-level positions through which IT professionals historically developed the experience required for senior roles are precisely the positions most subject to AI substitution, producing the “pipeline sustainability problem” documented in
This framework resolves the apparent paradox between rising senior salaries and falling total employment: the labor market is not contracting uniformly but reorganizing across all three stages simultaneously. The policy and educational implications flow directly from this structure. Curricula premised on creation-task proficiency are becoming misaligned not merely because AI can perform those tasks, but because the validation tasks that replace them require different cognitive skills—adversarial reasoning about AI failure modes, contextual judgment about output quality, and domain expertise sufficient to recognize hallucination—which current training pipelines do not systematically develop.
6 Implications
6.1 Policy implications
The empirical findings carry several implications for policy design, though the caution appropriate to a literature still in its early stages applies throughout.
Education systems face adjustment pressure that the available evidence documents as real but that outpaces the evidentiary base for prescriptive specifics.
The concentration of displacement among junior and entry-level roles has more direct policy implications, supported by convergent evidence across multiple high-quality studies.
The gender dimension of AI-driven displacement—with disproportionate exposure in female-dominated clerical and content creation roles—requires gender-responsive labor market policy design, including targeted retraining support and transition assistance.
The geographic evidence—showing potentially more severe disruption in developing economies despite lower AI adoption rates—suggests that internationally coordinated policy responses, including technology transfer and capability-building support, may be necessary to prevent growing divergence between AI-adopting and non-adopting economies.
6.2 Implications for theory
The empirical evidence reviewed here requires updating of several theoretical frameworks. Task-based models (
The induced innovation framework predicts that low-wage economies should experience slower AI adoption and therefore less displacement. However, ILO (
7 Limitations of this review
Several limitations of the present review must be acknowledged. First, the literature itself is nascent: the median study in our sample covers fewer than three years of post-ChatGPT data, limiting the ability to distinguish transient adjustment dynamics from permanent structural change. Second, publication bias may inflate effect sizes: studies documenting significant displacement effects are more likely to be submitted and published than null results. Third, geographic coverage is heavily skewed toward high-income, English-language contexts, limiting generalizability to the majority of the global labor force. Fourth, many key outcomes—particularly worker welfare, mental health, and the quality of AI-adjacent jobs that are growing—are not captured by the job posting and platform data that dominate the literature.
8 Future research directions
The synthesis points toward several priority research gaps. Worker-level longitudinal studies—tracking displaced individuals through the labor market transition, measuring reemployment rates, wage trajectories, and wellbeing outcomes—are essential to complement the employer-demand data that currently dominate the literature. Cross-national replication in developing and emerging economies, particularly in Southeast Asia, South Asia, Sub-Saharan Africa, and Latin America (where cognitive services outsourcing represents a significant employment base), is urgently needed.
Causal identification of the productivity-vs.-displacement channel requires firm-level data linking AI tool adoption, individual productivity measures, and employment decisions—data that are rarely available and require cooperation between researchers and employers. The Agent Era of autonomous AI systems (2025–present) represents a new treatment condition whose labor market effects have barely begun to be studied; continuous longitudinal monitoring of the platforms and job boards analyzed in prior studies is a methodological priority.
Finally, the non-employment dimensions of labor market adjustment—hours worked, task composition within surviving roles, the quality and stability of AI-adjacent positions, and informal labor market effects—require methodologies beyond job posting analysis, including survey instruments specifically designed to capture AI-mediated task change.
9 Conclusion
This systematic review assembles the first comprehensive synthesis of observed, post-LLM empirical evidence on AI-driven labor market displacement. The evidence base, though still limited by short observation windows and significant geographic gaps, already documents meaningful and statistically robust changes in job posting volumes, task demand, wage structures, and platform labor markets attributable to or consistent with LLM deployment.
The headline finding is that observable labor market effects are already present and exceed what would be expected from prior automation waves at comparable stages of adoption. A 14%–41% decline in postings for LLM-exposed roles across multiple geographies and platforms, a 15%–22% AI skill wage premium, and the near-elimination of certain task categories on online labor platforms collectively suggest that LLMs are not merely transforming cognitive work but partially displacing it—particularly at the junior and mid-level of the skill distribution.
At the same time, the evidence resists a simplistic displacement narrative. Validation, security, and infrastructure orchestration roles are growing. Salaries at the top of the distribution are rising. The most sophisticated human cognitive work—requiring contextual judgment, relational intelligence, and novel problem-solving—remains largely undisrupted in the evidence to date. The labor market is not disappearing; it is reorganizing around a new set of human comparative advantages in an AI-augmented production environment.
The urgency of the research agenda cannot be overstated. If the Agent Era inaugurated in 2025 by autonomous AI systems represents a further acceleration of the displacement dynamics documented here, the window for evidence-informed policy design is narrowing rapidly. This review provides the empirical foundation; the work of translation into policy and institutional response must proceed with commensurate speed.
Statements
Author contributions
ND: Writing – original draft, Writing – review & editing.
Funding
The author(s) declared that financial support was not received for this work and/or its publication.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
1
AcemogluD. (2024). The simple macroeconomics of AI. NBER Working Paper No. 32487.
2
AcemogluD.RestrepoP. (2018). The race between man and machine: implications of technology for growth, factor shares, and employment. Am. Econ. Rev.108 (6), 1488–1542. 10.1257/aer.20160696
3
AcemogluD.RestrepoP. (2020). Robots and jobs: evidence from US labor markets. J. Polit. Econ.128 (6), 2188–2244. 10.1086/705716
4
AMRO (2025). Can the Philippines IT-BPM Industry Stay Ahead Amid the AI Wave? ASEAN+3 Macroeconomic Research Office Analytical Note. Singapore: AMRO.
5
AutorD.ChinC.SalomonsA.SeegmillerB. (2024). New frontiers: the origins and content of new work, 1940–2018. Q. J. Econ.139 (3), 1399–1465. 10.1093/qje/qjae008
6
AutorD. H. (2015). Why are there still so many jobs? The history and future of workplace automation. J. Econ. Perspect.29 (3), 3–30. 10.1257/jep.29.3.3
7
AutorD. H.LevyF.MurnaneR. J. (2003). The skill content of recent technological change: an empirical exploration. Q. J. Econ.118 (4), 1279–1333. 10.1162/003355303322552801
8
BerkesE.Bruns-SmithD.GuvenenF.MenzioG. (2024). AI-augmented vs. AI-replaced: Differential employment effects of large language models across OECD occupations. IZA Discussion Paper 17081.
9
Brookings Institution (2025). Reimagining the Future of Data and AI Labor in the Global South. Washington, DC: Brookings.
10
BrynjolfssonE. (2022). The turing trap: the promise & peril of human-like artificial intelligence. Dædalus151 (2), 272–287. 10.1162/daed_a_01915
11
BrynjolfssonE.ChandarB.ChenR. (2025). Canaries in the Coal Mine? Six Facts About the Recent Employment Effects of Artificial Intelligence. Stanford Digital Economy Lab Working Paper. Stanford, CA: Stanford University.
12
BrynjolfssonE.LiD.RaymondL. R. (2023). Generative AI at work. NBER working paper No. 31161. Publ. Q. J. Econ.140 (2), 889–942. 10.1093/qje/qjae044
13
CazzanigaM.PizzinelliC.RockallE.TavaresM. M. (2024). Exposure to artificial intelligence and occupational mobility: A cross-country analysis. IMF Working Paper No. 24/116.
14
Cedefop (2024). AI and the Future of Work in Europe: Evidence from Online job Advertisements 2021–2024. Cedefop Research Paper No. 89. European Centre for the Development of Vocational Training.
15
Codemix Analytics (2024). The Developer Hiring Market 2022–2024: AI Coding Assistants and Entry-level job Dynamics. Codemix Analytics Industry Report, Q3 2024.
16
CNBC (2025). Why India's IT sector is shedding jobs. CNBC News, August 4. Available at: https://www.cnbc.com/2025/08/04/indias-it-layoffs-spark-fears-ai-is-hurting-jobs-in-critical-sector.html (Accessed April 29, 2026).
17
Dell'AcquaF.McFowlandE.MollickE. R.Lifshitz-AssafH.KelloggK.RajendranS.et al (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School Working Paper 24-013.
18
EloundouT.ManningS.MishkinP.RockD. (2024). GPTs are GPTs: labor market impact potential of LLMs. Science384 (6702), 1306–1308. 10.1126/science.adj0998
19
FreyC. B.OsborneM. A. (2017). The future of employment: how susceptible are jobs to computerisation?Technol. Forecast. Soc. Change.114, 254–280. 10.1016/j.techfore.2016.08.019
20
GoosM.ManningA.SalomonsA. (2014). Explaining job polarization: routine-biased technological change and offshoring. Am. Econ. Rev.104 (8), 2509–2526. 10.1257/aer.104.8.2509
21
GregoryT.SalomonsA.ZierahnU. (2021). Racing with or against the machine? Evidence from Europe. J. Eur. Econ. Assoc.20 (2), 869–899. 10.1093/jeea/jvab040
22
HuiX.ReshefO.ZhouL. (2024). The short-term effects of generative artificial intelligence on employment: evidence from an online labor market. Organ. Sci.35 (6), 1977–1989. 10.1287/orsc.2023.18441
23
IMF (2025). Artificial intelligence and the Philippine labor market. IMF Working Paper No. 25/043.
24
International Labour Organization (2023). Generative AI and Jobs: A Global Analysis of Potential Effects on job Quantity and Quality. ILO Working Paper 96. Geneva: International Labour Office.
25
IpeirotisP.PapadimitriouP.GoldsmithJ. (2024). The impact of large language models on crowdwork: evidence from Amazon mechanical Turk, 2019–2024. ACM Conference on Human Factors in Computing Systems (CHI 2024) Extended Abstracts.
26
KozlowskiM.LewandowskiP.RutkowskiJ. (2024). LLMS and the Polish IT Labor Market: Evidence from job Posting Panels 2018–2024. IBS Research Report 02/2024. Warsaw: Institute for Structural Research.
27
KunstS.KuhnP.MuellerA. (2024). AI Skill premiums in European labor markets: evidence from linked employer-employee data. Eur. Econ. Rev.168, 104823.
28
McKinsey Global Institute (2023). The Economic Potential of Generative AI: The Next Productivity Frontier. New York City, NY: McKinsey & Company.
29
MokyrJ.VickersC.ZiebarthN. L. (2015). The history of technological anxiety and the future of economic growth: is this time different?J. Econ. Perspect.29 (3), 31–50. 10.1257/jep.29.3.31
30
NASSCOMB. C. G. (2025). Roadmap for job Creation in the AI Economy. NITI Aayog Frontier Tech Hub Report. New Delhi: NITI Aayog.
31
NedelkoskaL.QuintiniG. (2018). Automation, skills use and training. OECD Social, Employment and Migration Working Papers, No. 202.
32
NoyS.ZhangW. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science381 (6654), 187–192. 10.1126/science.adh2586
33
ODI (2024). The AI Time Bomb: 2.5 Million Jobs at Risk—is Kenya Ready?London: Overseas Development Institute.
34
OkinyiM. (2024). “Impact of Remotasks Closure on Kenyan Workers”, in Data Workers‘ Inquiry, eds MiceliM.DinikaA.KauffmanK.Salim WagnerC.SachenbacherL. (Oakland, CA: DAIR Institute). https://data-workers.org/mophat
35
PageM. J.McKenzieJ. E.BossuytP. M.BoutronI.HoffmannT. C.MulrowC. D.et al (2021). The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. Br. Med. J.372, n71. 10.1136/bmj.n71
36
PerryN.SrivastavaM.KumarD.BonehD. (2023). Do users write more insecure code with AI assistants?Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, 2785–2799.
37
PizzinelliC.PantonA.TavaresM. M.CazzanigaM.LiL. (2023). Labor market exposure to AI: Cross-country differences and distributional implications. IMF Working Paper No. 23/216.
38
Rest of World (2024). Scale AI's Remotasks is booting workers with no explanation. Rest of World, April 2. Available at: https://restofworld.org/2024/scale-ai-remotasks-banned-workers/ (Accessed April 29, 2026).
39
WebbM. (2020). The impact of artificial intelligence on the labor market. SSRN Working Paper 3482150.
40
WilesJ.ConnerM.O'BrienD. (2024). Generative AI and Content Creation Employment: Lightcast Labor Market Analysis 2020–2024. Boston, MA: Lightcast Research Report.
Summary
Keywords
artificial intelligence, chatGPT, job postings, labor displacement, large language models, systematic review, technological unemployment, wage polarization
Citation
Dehouche N (2026) Creation, validation, obsolescence: observed evidence of AI-driven labor market displacement, 2020–2025. Front. Hum. Dyn. 8:1815037. doi: 10.3389/fhumd.2026.1815037
Received
21 February 2026
Revised
02 April 2026
Accepted
14 April 2026
Published
07 May 2026
Volume
8 - 2026
Edited by
Abdul Shaban, Tata Institute of Social Sciences, India
Reviewed by
Sandun Dassanayake, University of Moratuwa, Sri Lanka
Md Khaja Mohiddin, Bhilai Institute of Technology, Raipur, India
Updates

Check for updates
Copyright
© 2026 Dehouche.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Nassim Dehouche Nassim.deh@mahidol.edu
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.