Abstract
Malignant melanoma (MM) is the most aggressive form of skin cancer, for which early detection is critical and strongly associated with improved survival outcomes. Recent advances in large language models (LLMs), such as ChatGPT and Gemini, present promising opportunities to support melanoma early screening and clinical decision-making. However, despite increasing interest in LLM-based dermatologic applications, their diagnostic reliability across different populations remains insufficiently characterized. In this study, we systematically evaluated the performance of GPT-5.2 across skin pigmentation groups using Milk10K, a clinically curated, publicly available dermatology dataset comprising paired dermoscopic and clinical close-up images with histopathology-confirmed diagnoses and standardized skin tone annotations. GPT-5.2 was assessed on two clinically relevant tasks: binary malignancy discrimination and top-3 differential diagnosis. A balanced subset of 460 lesions (92 per skin tone class) was randomly selected for evaluation. Across both tasks and imaging conditions, GPT-5.2 showed moderate diagnostic performance, with broadly consistent accuracy, F1 score, and Cohen’s κ across skin tone groups, without evidence of systematic performance decline in darker skin tones. The incorporation of clinical close-up images provided modest improvements in overall performance while maintaining similar behavior across pigmentation classes. These findings suggest that GPT-5.2 exhibits stable melanoma-related diagnostic performance across diverse skin tones on this dataset. The study’s limitations and implications for future development are also discussed.
1 Introduction
Malignant melanoma (MM) is a malignancy of melanocytes, the pigment-producing cells of the skin. It is the most aggressive form of skin cancer and has a high propensity for metastasis to regional and distant organs. The American Cancer Society (ACS) projects approximately 112,000 new MM cases in 2026, including an estimated 65,400 cases in men and 46,600 in women (). An estimated 8,510 deaths are expected in the same year.
Melanoma comprises several clinically distinct subtypes that differ in growth patterns, anatomical distribution, and diagnostic visibility. Superficial spreading melanoma (SSM) is the most common subtype and often exhibits the classic asymmetry, border irregularity, color variation, diameter, and evolution (ABCDE) features, whereas nodular melanoma (NM) frequently grows rapidly and may lack these characteristic signs, complicating early recognition. Acral lentiginous melanoma (ALM), which arises on the palms, soles, or nail apparatus, represents a relatively uncommon subtype in populations of European ancestry but constitutes a larger proportion of melanoma cases in individuals with darker skin tones. Melanoma in situ (MIS), the earliest stage of disease confined to the epidermis, offers the greatest opportunity for curative treatment when detected promptly. These heterogeneous clinical presentations contribute to diagnostic complexity and highlight the need for improved detection strategies across melanoma subtypes ().
Early identification of melanoma is strongly associated with improved patient survival. However, diagnostic accuracy remains challenging because early melanoma lesions can resemble benign pigmented lesions such as nevi or seborrheic keratoses. These challenges may be further compounded by variations in skin pigmentation and lesion morphology across populations. Recent studies have therefore explored novel imaging and computational techniques to enhance melanoma detection. For example, Lin et al. have suggested that the spectrum-aided vision enhancer (SAVE) can improve the visualization of subtle pigmentary patterns and structural features in several melanoma subtypes, including ALM, MIS, NM, and SSM ().
The timely detection of melanoma is constrained by limited access to dermatologic care in many areas. In the United States, dermatology residency programs offer only 574 positions annually, contributing to workforce constraints. A recent study reported a mean wait time of 50 days for a new patient dermatology appointment (). These barriers to timely dermatologic evaluation may contribute to delayed detection, disease progression, and increased melanoma-related mortality rates.
The widespread adoption of large language models (LLMs) has the potential to help address persistent gaps in melanoma care. Recent LLMs extend beyond text processing and are now capable of analyzing, interpreting, and responding to user-uploaded medical images. Beginning with ChatGPT-4 Omni (GPT-4o), newer generations of LLMs have incorporated real-time multimodal capabilities, enabling the integration of textual, visual, and auditory inputs for complex diagnostic tasks (). The recently released GPT-5 further enhances diagnostic reasoning, reduces hallucinations, and improves robustness across heterogeneous clinical inputs (). Together, these advancements position LLMs as a promising frontier in medical artificial intelligence (AI) ().
Compared with traditional dermatologic resources, LLMs offer distinct advantages, including continuous availability, low cost, and scalable diagnostic support. Consequently, the public increasingly uses them for medication information, self-diagnosis, and disease-prevention guidance (). Clinicians and medical students also explore LLMs for knowledge acquisition and clinical decision support (, , ).
Numerous studies have investigated the use of LLMs for melanoma detection. Shifai et al. evaluated GPT-4 Vision (GPT-4 V) using 100 dermoscopic images randomly selected from the International Skin Imaging Collaboration (ISIC) Archive (). Similarly, Liu et al. analyzed 100 ISIC images to compare GPT-4 V with another LLM, Claude 3 Opus (). Sattler et al. assessed the performance of GPT-4 Turbo (GPT-4 T) and GPT-4o for melanoma identification using the Human Against Machine with 10,000 training images (HAM10K) dataset, reporting variable sensitivity, specificity, and accuracy (). In another study, Boostani et al. compared ChatGPT-4o with Gemini 2.0 Flash and found that model performance differed between clinical and dermoscopic images; notably, the combined use of both models enabled identification of melanoma in >90% of cases (). More recently, Wang et al. evaluated GPT-5 on images from both ISIC and HAM10K, demonstrating moderate performance in top-1 prediction and malignancy discrimination, along with substantially improved top-3 diagnostic accuracy ().
In addition to the ISIC and HAM10K, other datasets such as PH2 have also been used to evaluate LLMs (). However, these datasets predominantly contain images of light-skinned individuals and lack standardized skin tone annotations, limiting conclusions regarding model performance across all populations. Importantly, previous studies have shown that AI-assisted dermatologic diagnosis is less accurate for images of darker skin compared with lighter skin (), underscoring the need for rigorous evaluation of LLMs across diverse cutaneous phenotypes to ensure safe and equitable applications. Prior attempts to examine variation in skin tone have been constrained by small sample sizes (), and these assessments do not reflect the capabilities of newer models. To address this gap, we systematically evaluated the newly released GPT-5.2 across diverse skin complexions. To the best of our knowledge, this study represents the first systematic evaluation of GPT-5.2 for melanoma diagnosis across explicitly stratified skin tone groups using a dermatologic dataset with standardized skin tone annotations.
2 Methods
Despite rapid advances and growing interest in dermatologic AI, publicly available, expertly curated, and pathologically confirmed melanoma image datasets with diverse skin tones remain limited and difficult to access (, ). After surveying available dermoscopic image resources, we identified the Multimodal Imaging Learning Kit with 10,480 images (Milk10K) as the most suitable dataset for comparing LLM performance across skin tones (). Milk10K is a clinically curated dermatology benchmark dataset that provides histopathology-confirmed diagnoses and multimodal imaging, thereby offering a reliable resource for evaluating diagnostic algorithms across diverse populations. The Milk10K dataset comprises 5,240 image pairs consisting of dermoscopic and corresponding clinical close-up images, each accompanied by diagnostic labels and skin tone annotations ().
Skin pigmentation in Milk10K is annotated using an ordinal variable ranging from 0 (darkest) to 5 (lightest), inspired by the Fitzpatrick and Monk skin tone frameworks (26). For simplicity, we refer to this variable as MST (Milk10K Skin Tone) throughout this study.” No other changes are needed. Since the number of cases labeled as class 0 was very small (n = 6), this category was merged with class 1 prior to analysis. Milk10K additionally provides Fitzpatrick skin type annotations (I–VI) as auxiliary metadata; however, the MST-derived variable was used as the primary stratification variable in our analyses.
To enable a controlled evaluation of model performance across pigmentation groups, we constructed a balanced subset stratified by skin tone class. After merging MST classes 0 and 1, the combined MST 0/1 category contained 111 lesions. Lesions labeled as indeterminate epidermal proliferation represent diagnostically ambiguous entities that cannot reliably be categorized as benign or malignant and were therefore excluded from the analysis. Following this exclusion, the merged MST 0/1 group comprised 92 eligible lesions. We then randomly sampled 92 lesions from each of the remaining skin tone classes (MST 2–5) to achieve equal representation across groups. This stratified sampling strategy yielded a final evaluation cohort of 460 unique lesions (92 per skin tone group) and 920 corresponding images, which were used consistently across all experiments.
Given that major LLM-based benchmarks in this field have primarily evaluated ChatGPT models (, , ), this study focuses on GPT-5.2, a leading model released by OpenAI in December 2025, to facilitate comparability with prior analyses. Since previous studies have demonstrated that GPT-5 is not well-suited for top-1 diagnosis (), defined as the accuracy of the model’s single highest-ranked prediction, we did not include top-1 evaluation in the present investigation. Instead, GPT-5.2 was evaluated across two related diagnostic tasks: (A) top-3 differential diagnosis, defined as the ordered list of the three diagnoses ranked most likely by the model, and (B) malignancy discrimination, a binary classification of lesions as malignant or benign. For the top-3 task, a prediction was considered correct if the ground-truth diagnosis appeared anywhere among the model’s three highest-ranked outputs.
GPT-5.2 was accessed programmatically through the OpenAI Application Programming Interface (API). To reflect real-world deployment, the model was used “as is,” without fine-tuning or external training. Dermoscopic images were submitted to GPT-5.2 with standardized prompts, which consisted of two components: (1) an instruction specifying the diagnostic task and (2) a formatting instruction requesting output in JSON format for standardized downstream analysis. Model responses were parsed, stored, and compared with ground-truth clinical diagnoses using a custom analysis pipeline implemented in Python (version 3.11.10). The Python scripts developed for this study are publicly available at https://github.com/qwangmsk/Melanoma-Detect.
Model performance was evaluated using sensitivity, specificity, accuracy, F1 score, and related metrics, which were computed using R (version 4.3.3) and visualized using the ggplot2 package (version 3.5.1). To assess potential differences across skin tone groups, pairwise comparisons were performed using bootstrap resampling distributions. Two-sided p-values were derived from the empirical distributions of F1 score differences and were adjusted using the Benjamini–Hochberg procedure to control the false discovery rate.
3 Results
Using the evaluation cohort described in the Methods, GPT-5.2 predictions were generated for each lesion under both imaging conditions (dermoscopy-only and dermoscopy plus clinical close-up images) and across the two diagnostic tasks (top-3 differential diagnosis and binary malignancy discrimination). Model outputs were standardized and compared with ground-truth labels to compute performance metrics, both overall and stratified by skin tone class. These results were used to generate the summary tables and figures presented below, as well as Supplementary Tables S1–S3 and Supplementary Figure S1.
Table 1 and Figure 1 summarize GPT-5.2 performance for the top-3 differential diagnosis task across skin tone classes under both dermoscopy-only and combined dermoscopy plus clinical close-up settings. Across all evaluated metrics, including accuracy, recall, specificity, precision, F1 score, and Cohen’s kappa (κ), performance was broadly consistent across pigmentation groups, with no systematic decline observed in darker skin tones compared to lighter skin tones. Although modest variability was present among individual skin tone classes, the performance estimates overlapped substantially, and confidence intervals (CIs) exhibited considerable concordance, indicating the absence of a clear skin tone-dependent performance gradient. These findings suggest that GPT-5.2 prioritizes diagnostically relevant candidates with comparable effectiveness across diverse pigmentation levels.
Table 1
| (a) Dermoscopy only | |||||||
|---|---|---|---|---|---|---|---|
| Skin tone class | N | Accuracy | Recall | Specificity | Precision | F1 score | Kappa |
| 1 | 92 | 0.703 | 0.947 | 0.528 | 0.590 | 0.727 | 0.438 |
| 2 | 92 | 0.674 | 0.896 | 0.432 | 0.632 | 0.741 | 0.334 |
| 3 | 92 | 0.620 | 0.833 | 0.386 | 0.597 | 0.696 | 0.224 |
| 4 | 92 | 0.630 | 0.813 | 0.432 | 0.609 | 0.696 | 0.248 |
| 5 | 92 | 0.663 | 0.875 | 0.432 | 0.627 | 0.730 | 0.312 |
| Overall | 460 | 0.658 | 0.870 | 0.445 | 0.612 | 0.718 | 0.315 |
| (b) Dermoscopy + clinical close-up | |||||||
|---|---|---|---|---|---|---|---|
| Skin tone class | N | Accuracy | Recall | Specificity | Precision | F1 score | Kappa |
| 1 | 92 | 0.696 | 0.868 | 0.574 | 0.589 | 0.702 | 0.413 |
| 2 | 92 | 0.685 | 0.792 | 0.568 | 0.667 | 0.724 | 0.363 |
| 3 | 92 | 0.663 | 0.896 | 0.409 | 0.623 | 0.735 | 0.311 |
| 4 | 92 | 0.663 | 0.792 | 0.523 | 0.644 | 0.710 | 0.318 |
| 5 | 92 | 0.696 | 0.833 | 0.545 | 0.667 | 0.741 | 0.383 |
| Overall | 460 | 0.680 | 0.835 | 0.526 | 0.638 | 0.723 | 0.361 |
GPT-5.2 top-3 differential diagnosis performance across skin tones.
Figure 1
The addition of clinical close-up images resulted in modest overall improvements compared with dermoscopy alone (Table 1), particularly for specificity, precision, and Cohen’s κ, while recall remained relatively stable across skin tones. Importantly, these improvements were observed across all pigmentation groups without evidence of differential benefit by skin tone. Overall, GPT-5.2 exhibited stable performance for top-3 differential diagnoses across diverse skin tones on the Milk10K dataset, with no statistically significant disparities attributable to pigmentation level.
As shown in Figures 1B1,B2, the receiver operating characteristic (ROC) analyses were constructed using the confidence score associated with the model’s highest-ranked prediction for each lesion; therefore, they primarily reflect discrimination performance for the top-1 diagnosis rather than the full top-3 differential list. Notably, a recent study by Wang et al. also reported relatively lower performance for GPT-5 when evaluated using only the top-1 diagnosis (), which aligns with the ROC patterns observed for GPT-5.2 in the present study. The consistent pattern demonstrated in both the present study and Wang et al.’s study, in which top-1 diagnostic performance is lower than top-3 performance, provides additional validation that model accuracy improves when multiple differential diagnoses are considered.
Consistent with the findings for differential diagnosis prioritization, GPT-5.2 showed broadly comparable performance across skin tone classes in the binary malignancy discrimination task under both dermoscopy-only and combined dermoscopy plus clinical close-up conditions (Table 2; Figure 2). Across all evaluated metrics, no consistent pattern of performance decline was observed from lighter to darker skin tones. Although some variability was present among individual skin tone groups, pairwise comparisons based on bootstrap-derived F1 score distributions did not reveal statistically significant differences after multiple comparison adjustment. These findings are consistent with the substantial overlap in confidence intervals observed across skin tone groups, as shown in Figure 2.
Table 2
| (a) Dermoscopy only | |||||||
|---|---|---|---|---|---|---|---|
| Skin tone class | N | Accuracy | Recall | Specificity | Precision | F1 score | Kappa |
| 1 | 92 | 0.630 | 0.789 | 0.519 | 0.536 | 0.638 | 0.288 |
| 2 | 92 | 0.663 | 0.833 | 0.477 | 0.635 | 0.721 | 0.315 |
| 3 | 92 | 0.533 | 0.396 | 0.682 | 0.576 | 0.469 | 0.077 |
| 4 | 92 | 0.413 | 0.375 | 0.455 | 0.429 | 0.400 | −0.169 |
| 5 | 92 | 0.554 | 0.479 | 0.636 | 0.590 | 0.529 | 0.115 |
| Overall | 460 | 0.559 | 0.565 | 0.552 | 0.558 | 0.562 | 0.117 |
| (b) Dermoscopy + clinical close-up | |||||||
|---|---|---|---|---|---|---|---|
| Skin tone class | N | Accuracy | Recall | Specificity | Precision | F1 score | Kappa |
| 1 | 92 | 0.728 | 0.737 | 0.722 | 0.651 | 0.691 | 0.450 |
| 2 | 92 | 0.663 | 0.729 | 0.591 | 0.660 | 0.693 | 0.322 |
| 3 | 92 | 0.576 | 0.500 | 0.659 | 0.615 | 0.552 | 0.158 |
| 4 | 92 | 0.522 | 0.458 | 0.591 | 0.550 | 0.500 | 0.049 |
| 5 | 92 | 0.554 | 0.521 | 0.591 | 0.581 | 0.549 | 0.111 |
| Overall | 460 | 0.609 | 0.583 | 0.635 | 0.615 | 0.598 | 0.217 |
GPT-5.2 malignancy discrimination performance across skin tones.
Figure 2
The incorporation of paired clinical close-up images resulted in modest improvements in overall performance metrics; however, these gains were observed across multiple skin tone groups rather than being confined to any specific pigmentation category. ROC curves similarly demonstrated comparable discrimination ability across skin tones in both imaging scenarios (Figure 2). Collectively, these findings suggest that GPT-5.2’s malignancy discrimination performance on the Milk10K dataset is largely consistent across diverse skin pigmentation levels, with no evidence of significant performance disparity between darker and lighter skin tones. Taken together with the top-3 diagnosis results, the overall evidence suggests that GPT-5.2 maintains consistency in melanoma-related diagnostic performance across skin tone classes on this dataset. To provide additional clinical context for the malignancy discrimination task, Supplementary Table S3 compares GPT-5.2 with published clinician benchmarks for melanoma diagnosis. These contextual comparisons indicate that GPT-5.2 remains below the performance range reported for dermatologists.
Compared with the prior study evaluating GPT-5 on the dermoscopic datasets ISIC Archive and HAM10K (), GPT-5.2 demonstrated comparable but generally slightly lower performance on the Milk10K dataset (see Supplementary Table S1; Supplementary Figure S1 for direct comparison). In Wang et al.’s study (), GPT-5 achieved higher diagnostic metrics, whereas in the present study, GPT-5.2 reached top-3 diagnostic accuracy of approximately 0.66–0.68 with F1 scores of approximately 0.72, and malignancy discrimination accuracy of 0.56–0.61 with F1 scores of 0.56–0.60. These differences likely reflect the greater clinical heterogeneity and multimodal complexity of Milk10K compared with the ISIC Archive and HAM10K, rather than intrinsic model limitations. Importantly, a consistent pattern was observed in both the present study and Wang et al.’s study, in which GPT’s top-3 diagnostic performance exceeds malignancy discrimination performance, further supporting the validity of the present findings across datasets and model versions.
4 Discussion
To the best of our knowledge, this study represents one of the first systematic evaluations of a GPT-5.2 model across explicitly stratified skin tone groups using a large, pathologically confirmed dermatology dataset. Prior benchmarking studies of LLMs for melanoma detection have primarily relied on private datasets or public resources such as ISIC or HAM10K, which contain predominantly lighter skin types and lack standardized pigmentation annotations, thereby limiting conclusions regarding algorithmic equity. By leveraging the Milk10K dataset with skin tone-derived labels and balanced sampling across pigmentation classes, our study directly addresses this critical gap. Importantly, the absence of statistically meaningful performance differences across skin tones suggests that GPT-5.2 does not exhibit systematic degradation in diagnostic performance in darker skin within this dataset, in contrast to many AI-assisted systems that underperform in darker skin populations.
While GPT-5.2 exhibited moderate overall diagnostic accuracy, the stability of performance across pigmentation groups indicates that recent advances in LLM architecture may help mitigate some of the sources of skin tone-related bias reported in earlier computer vision systems. However, the performance variability across individual lesions and the only moderate agreement with ground truth underscore that the model is not suited to autonomous diagnosis. Accordingly, these findings should be interpreted as evidence of relative equity in model behavior rather than clinical readiness.
Several limitations of this study should be noted. First, although Milk10K provides explicit skin tone annotations and multimodal imaging, the dataset remains a curated research resource and may not fully capture the complexity of real-world clinical presentations. Consequently, the generalizability of these findings to routine clinical settings remains unclear. Second, the evaluation was performed using a balanced subset of lesions stratified by skin tone. This sampling strategy was intentionally adopted to minimize confounding due to unequal skin tone representation in the original dataset. However, since the resulting cohort does not reflect the true prevalence of melanoma in clinical populations, the reported performance metrics should be interpreted as comparative estimates under controlled conditions rather than direct estimates of clinical screening performance. Third, model performance was assessed using static images without additional clinical context (e.g., patient history, lesion evolution, or dermoscopic metadata) that clinicians typically incorporate into diagnostic reasoning, and the evaluation was conducted retrospectively on a single dataset due to the limited availability of public datasets with skin tone annotations. To support transparency and external validation, the identifiers of the 460 lesions included in this evaluation have been publicly released through our GitHub repository, enabling other investigators to reproduce our analysis or benchmark alternative models on the same cases. Finally, we did not assess calibration, decision thresholds, or user interaction effects, all of which are important considerations for clinical deployment. Future studies incorporating prospective validation, diverse patient populations, and human–AI interaction paradigms are needed to fully characterize the safety, reliability, and equity of LLM-based dermatologic decision support tools.
Given its moderate diagnostic performance and variability across cases, GPT-5.2 is best positioned as a clinician-supervised decision support tool rather than a standalone diagnostic system. Potential use cases include triage support, where the model may assist in identifying lesions that warrant expedited specialist evaluation, and educational support for trainees or patients seeking preliminary information. Importantly, the risk of false-negative predictions underscores the need for conservative triage strategies and clinician verification, particularly for lesions with ambiguous or high-risk features. In addition, practical considerations such as user interaction design, patient engagement, and medico-legal responsibility must be carefully addressed before clinical deployment. Accordingly, such systems should be considered adjunctive tools that complement, rather than replace, clinician expertise.
Statements
Data availability statement
Publicly available datasets were analyzed in this study. This data can be found at: https://api.isic-archive.com/doi/milk10k/.
Ethics statement
Ethical approval was not required for the study involving humans in accordance with the local legislation and institutional requirements. Written informed consent to participate in this study was not required from the participants or the participants’ legal guardians/next of kin in accordance with the national legislation and the institutional requirements.
Author contributions
KF: Data curation, Investigation, Methodology, Writing – original draft, Writing – review & editing. SA: Funding acquisition, Project administration, Resources, Supervision, Writing – review & editing. QW: Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing.
Funding
The author(s) declared that financial support was received for this work and/or its publication. This research was funded by Meharry’s American Cancer Society (ACS) (grant no. DICRIDG-21-071-01-DICRIDG) and the National Institute of Minority Health Disparities (NIMHD) (U54MD007586) and the National Institute of General Medical Sciences (NIGMS) (R16GM149359).
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that Generative AI was used in the creation of this manuscript. During the preparation of this manuscript, the authors used GPT-5.2 to assist with minor editing, language polishing, and drafting of computational scripts for data analysis and visualization. All outputs were independently verified, revised, and approved by the authors. The authors take full responsibility for the integrity of the data, the accuracy of the analyses, and the final version of the manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Author disclaimer
The views and conclusions contained in this paper are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the ACS, NIH, and Meharry Medical College.
Supplementary material
The Supplementary material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fmed.2026.1816102/full#supplementary-material
References
1.
Key Statistics for Melanoma Skin CancerCancer.org: American Cancer Society; (2026). Available online at: https://www.cancer.org/cancer/types/melanoma-skin-cancer/about/key-statistics.html (Accessed February 20, 2026).
2.
LuC-TLinT-LMukundanAKarmakarRChandrasekarAChangW-Yet al. Skin Cancer: epidemiology, screening and clinical features of Acral lentiginous melanoma (ALM), melanoma in situ (MIS), nodular melanoma (NM) and superficial spreading melanoma (SSM). J Cancer. (2025) 16:3972–90. doi: 10.7150/jca.116362,
3.
LinT-LKarmakarRMukundanAChaudhariSHsiaoY-PHsiehS-Cet al. Assessing the efficacy of the Spectrum-aided vision enhancer (SAVE) to detect Acral lentiginous melanoma, melanoma in situ, nodular melanoma, and superficial spreading melanoma: part II. Diagnostics. (2025) 15:714. doi: 10.3390/diagnostics15060714,
4.
BaschCHHillyerGCGoldBBaschCE. Wait times for scheduling appointments with hospital affiliated dermatologists in new York City. Arch Dermatol Res. (2024) 316:530. doi: 10.1007/s00403-024-03249-w,
5.
OpenAI. GPT-5 System Card. (2025). Available online at: https://cdn.openai.com/gpt-5-system-card.pdf (Accessed February 20, 2026).
6.
Team C. Chameleon: mixed-modal early-fusion foundation models. arXiv. (2024)
7.
Gemini. Google; (2025). Available online at: https://gemini.google.com/app (Accessed February 20, 2026).
8.
FredericksonKLGuiHBarbieriJSDaneshjouR. Artificial intelligence use in acne diagnosis and management—a scoping review. Int J Dermatol. (2025) 65:437–43. doi: 10.1111/ijd.70110
9.
ShahsavarYChoudhuryA. User intentions to use ChatGPT for self-diagnosis and health-related purposes: cross-sectional survey study. JMIR Hum Factors. (2023) 10:e47564. doi: 10.2196/47564,
10.
Marley PresiadoAMLopesL.HamelLiz. KFF Health Misinformation Tracking Poll: Artificial Intelligence and Health Information. Available online at: : The Henry J. Kaiser Family Foundationhttps://www.kff.org; (2024). Available online at: https://www.kff.org/health-information-and-trust/poll-finding/kff-health-misinformation-tracking-poll-artificial-intelligence-and-health-information/ (Accessed February 20, 2026).
11.
GliadkovskayaA.Some Doctors are Using Public AI Chatbots like ChatGPT in Clinical Decisions. Available online at: https://www.fiercehealthcare.com: Fierce Healthcare; (2024) (Accessed February 20, 2026).
12.
KuroiwaTSarconAIbaraTYamadaEYamamotoATsukamotoKet al. The potential of ChatGPT as a self-diagnostic tool in common orthopedic diseases: exploratory study. J Med Internet Res. (2023) 25:e47621. doi: 10.2196/47621,
13.
DuDPaluchRStevensGMüllerC. Exploring patient trust in clinical advice from AI-driven LLMs like ChatGPT for self-diagnosis. arXiv. (2024)
14.
KisvardaySYanAYarahuanJKatsDJRayMKimEet al. ChatGPT use among pediatric health care providers: cross-sectional survey study. JMIR Forma Res. (2024) 8:e56797. doi: 10.2196/56797,
15.
OzkanETekinAOzkanMCCabreraDNivenADongY. Global Health care professionals’ perceptions of large language model use in practice: cross-sectional survey study. JMIR Med Educ. (2025) 11:e58801-e. doi: 10.2196/58801
16.
ShifaiNVan DoornRMalvehyJSangersTE. Can ChatGPT vision diagnose melanoma? An exploratory diagnostic accuracy study. J Am Acad Dermatol. (2024) 90:1057–9. doi: 10.1016/j.jaad.2023.12.062,
17.
LiuXDuanCKimM-KZhangLJeeEMaharjanBet al. Claude 3 opus and ChatGPT with GPT-4 in Dermoscopic image analysis for melanoma diagnosis: comparative performance analysis. JMIR Med Inform. (2024) 12:e59273. doi: 10.2196/59273,
18.
SattlerSSChetlaNChenMHageTRChangJGuoWYet al. Evaluating the Diagnostic accuracy of ChatGPT-4 Omni and ChatGPT-4 Turbo in identifying melanoma: comparative study. JMIR Dermatol (2025);8:e67551-e, doi: 10.2196/67551.
19.
BoostaniMLallasAGoldustMNádudvariNLőrinczKBánvölgyiAet al. Diagnostic performance of multimodal large language models in distinguishing melanoma from nevi in clinical and dermoscopic images. JAAD International. (2025) 23:58–60. doi: 10.1016/j.jdin.2025.08.008,
20.
WangQAmugoIRajakarunaHIrudayamMJXieHShankerAet al. Evaluating GPT-5 for melanoma detection using Dermoscopic images. Diagnostics. (2025) 15:3052. doi: 10.3390/diagnostics15233052,
21.
PerlmutterJWMilkovichJFremontSDattaSMosaA. Beyond the surface: assessing GPT-4's accuracy in detecting melanoma and suspicious skin lesions from Dermoscopic images. Plastic Surgery. (2025) 34:293–300. doi: 10.1177/22925503251315489
22.
GrohMBadriODaneshjouRKoochekAHarrisCSoenksenLRet al. Deep learning-aided decision support for diagnosis of skin disease across skin tones. Nat Med. (2024) 30:573–83. doi: 10.1038/s41591-023-02728-3
23.
CironeKAkroutMAbidLOakleyA. Assessing the utility of multimodal large language models (GPT-4 vision and large language and vision assistant) in identifying melanoma across different skin tones. JMIR Dermatology. (2024) 7:e55508. doi: 10.2196/55508,
24.
DaneshjouRVodrahalliKNovoaRAJenkinsMLiangWRotembergVet al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci Adv. (2022) 8:1–7. doi: 10.1126/sciadv.abq6147
25.
TschandlPAkayBNRosendahlCRotembergVTodorovskaVWeberJet al. MILK10k: a hierarchical multimodal imaging-learning toolkit for diagnosing pigmented and nonpigmented skin Cancer and its simulators. J Invest Dermatol. (2025) 146:357–364.e7. doi: 10.1016/j.jid.2025.06.1594
26.
MehtaHSarkarR. The monk skin tone scale: a tool dermatology should not overlook. J Am Acad Dermatol. (2025) 93:e139–41. doi: 10.1016/j.jaad.2025.05.1437,
Summary
Keywords
ChatGPT, dermoscopy, GPT-5.2, large language model, melanoma diagnosis, skintone
Citation
Frederickson KL, Adunyah SE and Wang Q (2026) Evaluation of GPT-5.2 for melanoma detection across skin tones. Front. Med. 13:1816102. doi: 10.3389/fmed.2026.1816102
Received
23 February 2026
Revised
20 March 2026
Accepted
23 April 2026
Published
08 May 2026
Volume
13 - 2026
Edited by
Gerardo Cazzato, University of Bari Aldo Moro, Italy
Reviewed by
Riya Karmakar, National Chung Cheng University, Taiwan
Michał Strzelecki, Lodz University of Technology, Poland
Updates
Copyright
© 2026 Frederickson, Adunyah and Wang.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Qingguo Wang, qiwang@mmc.edu
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.