ORIGINAL RESEARCH article

Front. Psychol., 23 September 2025

Sec. Cognition

Volume 16 - 2025 | https://doi.org/10.3389/fpsyg.2025.1640417

cpCST: a new continuous performance test for high-precision assessment of attention across the lifespan

  • 1. Brain Aging and Mental Health Laboratory, Clinical Research, Nathan Kline Institute, Orangeburg, NY, United States

  • 2. Design Acquisition and Neuromodulation Laboratories, Center for Biomedical Imaging and Neuromodulation, Nathan Kline Institute, Orangeburg, NY, United States

  • 3. Child Mind Institute, New York, NY, United States

  • 4. Columbia MR Research Center, Columbia University, New York, NY, United States

  • 5. Department of Psychiatry, New York University Medical Center, New York, NY, United States

Abstract

Introduction:

Assessing sustained attention presents methodological challenges, particularly when spanning diverse populations whose baseline sensorimotor functioning may vary significantly.

Methods:

This study introduces the Continuous Performance Critical Stability Task (cpCST), a novel paradigm combining high-density sampling of behavior (30 Hz), individualized calibration, and fixed-difficulty assessment to measure attentional control. In a sample of 166 adults (ages 18–76), we evaluated the psychometric properties of the cpCST’s instantaneous reaction time (iRT) metric derived through dynamic time warping.

Results:

The cpCST demonstrated exceptional reliability (bootstrap split-half r = 0.999) and predictive validity for cognitive performance (flanker and Woodcock-Johnson) and cardiorespiratory fitness (VO2submax). The task achieved high temporal efficiency, with just 2 min of data correlating at r = 0.94 with full-task performance, outperforming a standard arrow-based flanker task. The cpCST’s individualized calibration effectively isolated attentional control processes from baseline sensorimotor function, eliminating age-related slowing effects typically observed in reaction time tasks.

Discussion:

This approach offers methodological advantages for lifespan studies, clinical populations, integration with neurophysiological measures, and computational modeling approaches while addressing limitations of existing attention assessment paradigms.

1 Introduction

Attention is an intuitive concept that is considered a core component of cognition and everyday functioning (; ; ; ), exhibiting a clear trajectory of early life maturation and later-life decline (; ). Physiological factors such as fitness are well-known to have a general influence on cognition () and are thought to impact performance through improvements in attentional control (, ; ). Across the lifespan, attentional processes are linked to the successful navigation of a host of everyday behaviors (; ; ; ; ; ; ); and like many apparently simple behaviors, it can be challenging to define and measure (; ; ; ; ).

Experimental and clinical work focused on the measurement of sustained attention has produced a wide selection of continuous performance tasks (e.g., Continuous Performance Test: CPT (); AX-CPT (); Psychomotor Vigilance Test: PVT (); Paced Auditory Serial Attention Test: PASAT (); Test of Variables of Attention: TOVA (); Mackworth Clock Test (); Sustained Attention to Response Test: SART (); Cambridge Neuropsychological Test Automated Battery: CANTAB (); Continuous Visual Attention Test: CVAT (); Gradual-onset Continuous Performance Test: GradCPT (). These tasks are tuned to capture behavioral features thought to contribute to successful performance or identify specific areas of deficit, based on the paradigm and study population of interest (; ; ; ; ). Given the complexity of attentional processes and limitations inherent in any one particular paradigm, the development of a large corpus of measures with different features and tuning will facilitate continued knowledge building and translation to practical applications.

As outlined below, opportunities exist to augment or improve upon existing paradigms through novel behavioral sampling, dynamically adaptive assessment, and task calibration approaches - amongst others. Here we describe the Continuous Performance Critical Stability Task (cpCST), which modifies an established sensorimotor integration task to create a novel attention task featuring high-density behavioral sampling, dynamic adaptation, and effective behavioral calibration across the lifespan. We first describe these new features, not readily available in current paradigms, and the proposed advantages of these enhancements. We then present a preliminary psychometric evaluation of the cpCST’s primary outcome metric (instantaneous reaction time; iRT). We also examine predictive validity of the cpCST iRT to flanker task performance, Woodcock-Johnson Intellectual Ability and Achievement scores, as well as a measure of cardiorespiratory fitness (VO2max), before discussing the advantages of the cpCST paradigm in relating physiological and brain timeseries to participant behavior.

In most continuous performance tests, attention is probed via button press responses at discrete intervals ranging from roughly one to several (10+) seconds apart c.f. (; ; ); attentional lapses are inferred on the occasion of delayed, missed, or incorrect responses. Despite the relatively sparse sampling of behavior (< 1 Hz - once every second or longer), these response time studies have demonstrated attentional fluctuations over time (; ; ; ; ; ); however, higher density sampling of behavior may more effectively characterize the maintenance of focus over time, moment-to-moment fluctuations, and/or lapses in attention. While many established tasks require continuous monitoring of stimuli (e.g., CPT and PVT), they do not sample behavior continuously. Our primary goal for the development of the cpCST was to create a task that sampled behavior at a much higher rate (30 times per second) than existing tests. Further, unlike a stop-signal () or gradual-onset continuous performance test (), the cpCST does not include the feature of building up a prepotent response as the result of higher frequency responding. This allows for the assessment of continuous attention under qualitatively different conditions than many tests with higher responding rates.

Administering reaction time tasks to older adults, children, or clinical populations often requires adjustments to various parameters such as stimulus type, stimulus modality (audio vs. visual), presentation and response durations, response interval, interstimulus intervals, stimulus set sizes, or proportion of trial types (see ; ; ; ; ). These adjustments are motivated by group differences in sensorimotor speed, working memory, auditory or visual acuity, etc., (; ; ; ; ; ; ; ; ). While accommodations such as these allow for versions of standard neurocognitive tasks to be applied across a wider range of individuals, they raise concerns regarding the comparability of results across test variants (; ; ; ). Additionally, these changes are applied under the assumption that the altered parameters are uniformly appropriate to the group in question [e.g., trading arrow shapes for cartoon fishes in the ANT-C task or an increase in presentation duration for older adults (; )], despite well-documented heterogeneity within groups (; ; ). As part of the NKI-RS2 lifespan characterization study, our goal was to develop a test that did not require different versions across the ages 9–75 years. Our focus was to use simple stimuli and an intuitive response modality to decrease instructional or proficiency barriers. Further we adopted a closed-loop system developed for a sensorimotor paradigm (), described in detail below, that is calibrated to the individual’s own motor performance to equate individual performance differences into a uniform task design. We are not aware of any other continuous performance tasks that incorporate this design feature.

Rather than assuming that a single set of task adaptations will be equally appropriate across a given group (e.g., older adults or children), fully adaptive paradigms individualize task parameters for each participant by dynamically altering key task features in response to ongoing task behavior. Some tasks are explicitly designed to be adaptive [e.g., Stop-Signal Reaction Time ()]. More recent approaches overlay adaptive procedures that alter task features such as presentation time, response windows, or set size, in response to participant performance in real time as the task evolves (; , ; ). As such, each participant’s task is custom tailored to their individual performance on that task through approaches such as staircase or Bayesian-based adaptive algorithms (; ). These approaches are more efficient (; ; ; ), and can be leveraged not only in assessment, but also training protocols (e.g., ). However, they also suffer from drawbacks such as edge case and small sample size failures, induction of artifactual oscillatory “yo-yo” patterns in difficulty, as well as the additional complexity involved in dynamically adapting task parameters in real time (; ; ).

One promising approach is to leverage the best of both dynamic adaptive approaches and fixed stable approaches. Participant ability is assessed via adaptive staircase or Bayesian methods in a calibration phase. During the test phase, difficulty is set to a fixed level matching the participant’s individual ability (e.g., the stimulus onset asynchrony that resulted in > 70% accuracy) (; ; ). Under this approach, difficulty is individually tuned - thus avoiding assumptions about the appropriateness or comparability of group-specific stimulus changes. And the difficulty is fixed during the testing phase, reducing computational complexity and artifactual issues such as induced oscillatory behavior or algorithmic failure. For the cpCST, we leveraged such a hybrid approach by incorporating a calibration phase and a test phase. The goal was to maximize our ability to individually calibrate performance on the sensorimotor component of the task so that performance adjustments related to drifting attention could be better compared across individuals who differ in performance on the sensorimotor integration dimension of the task.

The cpCST design is intended to capture attentional dynamics from moment to moment using a simple, relatively short duration, high information density, individualized difficulty approach in order to maximize the detection of attentional control performance differences across a lifespan sample. The cpCST uses the hybrid-calibrated approach to first assess visuomotor ability, replicating the Critical Stability Task developed by , and then employs a fixed difficulty visuomotor continuous performance task based on the individual’s motor stability threshold (MST; see below). Thus the cpCST retains Critical Stability Task name and adds a new continuous performance phase. The cpCST additionally possesses useful features such as simple task instruction and continuous sampling of behavior (@30 Hz; i.e., 30 times per second), providing a robust complement to existing neurocognitive tools.

The cpCST is an extension of a psychomotor tracking task (Critical Stability Task; CST) developed for NASA by to evaluate pilot performance under unstable control conditions. This simple and elegant foundational paradigm has subsequently been adopted by human and non-human primate laboratories to develop and refine human-machine interfaces (), understand the neural mechanisms of sensorimotor coordination (), and examine drug induced motor control disruption (). We propose that the CST also provides a strong foundation upon which to build a continuous self-calibrated task to assess dynamic fluctuations in attentional control. Specifically, we employed a variant of the original CST to serve as a calibration phase that established a participant’s individual MST. We then used that individualized score to set the difficulty level for that participant’s continuous performance phase. The goal for this new continuous performance phase was to maintain attention on a relatively easy task (set as 30% of MST, based on internal pilot testing) for a 10 min period to capture attentional drift and the latency to respond to a drifting stimulus. It may be useful to clarify that there are a number of continuous tracking tasks in which participants use devices such as joysticks, trackballs, etc., to align a cursor with a spatially moving target item (; ; ; ; ). The main difference between these sorts of continuous tasks and the cpCST is that the cpCST is based on closed-loop paradigm in which the goal is to maintain the target stimulus at the center of the screen, rather than track, for example, a vertically oscillating target. Additionally, in both the task and the cpCST, the target’s stability is solely dependent on the user’s movements – there is no externally driven movement of a target object for the user to track. See Figure 1, described in more detail below.

FIGURE 1

, and more recently, . The participant (A) is tasked with keeping the stimulus disk at the center of the screen. They controlled the position of the central dot using a handheld controller enabled with an inertial measurement unit (B). To control the movement of the central stimulus disk, participants were required to counteract the drift of the central stimulus by tilting the inertial motion unit (IMU)-enable controller in the opposite direction. The interaction between the stimulus disk movements and participant movements drove an unstable system governed by the equation in box (C), where x(t) represents the horizontal position of the central stimulus disk at each time point, u(t) is the corresponding horizontal position of the participant’s (invisible) cursor. Lambda operates as a gain mechanism on the system, controlling the magnitude with which the discrepancy between participant and stimulus disk positions impact cursor position. Participants are provided visual feedback regarding the current position of the stimulus disk on the computer screen (D).

To evaluate this novel task, we embedded the cpCST within the Nathan Kline Institute Rockland Sample II (NKI-RS2), a large-scale, community-based lifespan study. The NKI-RS2 was designed to support the development and validation of next-generation tools for phenotyping normative brain-behavioral associations and investigate the underlying neural and physiological mechanisms that promote mental health across the lifespan. This context offered an opportunity to examine individual differences in attentional control across a wide age range using a task that prioritizes continuous behavioral sampling, individualized calibration, and high-density data collection. We characterized cpCST performance in relation to broader indices of cognitive function and health. In this preliminary analysis, we report behavioral data from the cpCST from a subset of participants to describe the development of a key task performance metric (instantaneous reaction time; iRT) and establish its reliability and preliminary predictive validity on cognitive and physiological indices.

2 Materials and methods

2.1 Participants

Participants were recruited into the Nathan Kline Institute Rockland Sample II study (NKI-RS2) through prior participation in the NKI-RS research program (), community outreach, and word-of-mouth. The lifespan sample recruited participants from age 9 to 76 years who were residents of Rockland, Orange, Bergen, or Westchester counties in the north suburban New York City area. All were fluent in English and had no severe physical or sensory limitations, contraindications for MRI or cardiovascular fitness testing, or acute psychiatric symptoms. Participants were excluded if they had a history of schizophrenia, schizoaffective disorder, autism spectrum disorder, or serious neurological conditions (e.g., Parkinson’s disease, traumatic brain injury, dementia). Current psychotropic medication use and serious medical conditions or metabolic disorders affecting the central nervous system (e.g., malignancy, HIV) were also exclusionary. For this preliminary analysis, we included a subset of 166 participants aged 18 to 76 years (M = 51.61, SD = 16.36), 66% female, with complete and quality controlled data for the cpCST and cardiorespiratory fitness procedure. Please note, this analysis is based on a convenience sample extracted from the ongoing NKI-RS2 characterization study to present preliminary findings and introduce novel task development.

2.2 Procedures

Sample characterization data were collected via remotely administered surveys on the MindLogger Platform () and in-person testing. Demographic data was collected via Mindlogger surveys, clinical characterizations were conducted by research staff in-person and via telephone interviews; all cognitive and cardiorespiratory fitness data were collected onsite. The study was approved by the NKI Institutional Review Board, and all participants provided informed consent before undergoing any procedures.

2.3 Measures

2.3.1 Continuous Performance Critical Stability Task (cpCST)

The Continuous Performance Critical Stability Task (cpCST) was administered in a dedicated testing room at the Center for Biomedical Imaging and Neuromodulation (CBIN) at NKI. Participants were seated in front of a 61 × 36 cm computer monitor, at a distance of 65 cm. The monitor displayed a circular stimulus at the center of the screen subtending 3.17 degrees of visual angle (DVA). Screen resolution was 1,920 × 1,080 pixels. They were instructed to maintain the position of a circular stimulus at the center of the screen. The stimulus could move along one dimension (left-right on the x-axis). Participants controlled the position of the central stimulus by tilting a custom-built handheld inertial motion unit (IMU) that measured rotation along the x-axis to the left or right. Participants were given a brief (∼2 min) practice round in which they gained familiarity with the controls at a very low difficulty level prior to beginning the calibration phase.

During calibration, the gain parameter was linearly increased over time so that even small corrections by the participant resulted in large changes in the stimulus position – thus systematically increasing difficulty. The gain of the system was characterized by a lambda (λ) parameter (). If the participant failed to maintain the stimulus within a predefined spatial boundary (80% of distance from the center, or ± 22.28 DVA from center), it resulted in a “crash” and the circular stimulus was reset to the center of the screen. The position of the central stimulus and the user’s tracking position were continuously recorded at a sampling rate of 30 Hz. See Figure 1 for a schematic and equation describing the closed-loop unstable system that provides the dynamic conditions under which the participant must continuously provide corrective adjustments to stabilize the stimulus. See also Figure 2A for the cpCST task screen schematic.

FIGURE 2

The Continuous Performance Critical Stability Task (cpCST) employed a hybrid approach consisting of two distinct phases: an initial calibration phase that estimated the participant’s motor stability threshold, which was followed by a continuous performance phase in which they performed the task at a fixed difficulty level.

2.3.1.1 Calibration phase

The calibration phase was similar to the original approach. Specifically, we employed a maximal performance to failure protocol similar to working memory tasks like digit span, and Corsi blocks (; ), and conceptually similar to the testing-the-limits approach ().

During the calibration phase, participants attempted to maintain the central position of the stimulus by adaptively tilting the accelerometer device. Task difficulty (lambda) was linearly increased over time until the participant failed to control the stimulus - defined by the stimulus exceeding 80% of the distance from the center of the screen (crashed). See Figure 2B.

Following failure, the stimulus was reset to the center of the screen, and the lambda parameter was reset to 50% of the value achieved at the time of the crash - allowing participants to “reset” and build back up to a higher difficulty. This process was repeated 10 times. We estimated each participant’s overall motor stability threshold (MST) by calculating the average lambda values reached over the final three calibration trials. Given the fixed number of calibration trials, we assessed whether the calibration phase was effective in reaching a stable estimate of each participant’s MST by calculating the amount of time required for each participant to reach asymptotic performance. Over 95% of participants reached asymptote within 1.5 min. Only two participants failed to reach asymptote by the final trial. We did not remove participants from continuous performance analyses based on this calibration metric. See Figure 3.

FIGURE 3

2.3.1.2 Continuous performance phase

In the continuous performance phase, participants performed the same task as in the calibration phase. However, in this phase the difficulty level was held to just 30% of the participant’s individually estimated MST, and the trial duration was fixed at 10 min. See Figure 2C. To evaluate task compliance, we examined the participants’ mean position. Participants were able to maintain the position of the stimulus near the center of the screen, with an average distance of −0.156 ± 0.168 DVA across all participants. See Figure 2D.

2.3.2 Flanker task

The flanker task was administered in a dedicated testing room in the CBIN at NKI. Participants performed a modified version of a flanker paradigm (; ) in which they were asked to respond to a central target flanked by an array of distractors. Each trial presented one of three trial types: congruent, where the flanking stimuli matched the central target (e.g.,<<<<<); incongruent, where the flankers opposed the central target (e.g., <<><<); or neutral, where the flankers provided no directional information (e.g., - -<- -). Trial types were presented in equal proportions and were first-order counterbalanced to control for sequential effects.

The task consisted of a practice block of 30 trials with feedback. Each trial began with a fixation cross displayed for 500 ms, followed by the target stimulus. The inter-trial interval (ITI) averaged 1.16 s; mean total task duration was 618 s. Participants responded using a standard keyboard, pressing the “C” or “M” keys to indicate left or right central arrow directions, respectively. They were instructed to respond as quickly and accurately as possible. Participants were required to achieve at least 80% accuracy in the practice block to move on to the test phase. The test phase consisted of three blocks of 120 trials each, for a total of 360 trials. Participants were required to achieve at least 80% accuracy across all test trials to be included in analyses. Nine participants did not meet this minimum criteria.

2.3.3 Woodcock Johnson Tests of Cognitive Abilities and Tests of Achievement (WJ)

Participants were administered a subset of the Woodcock Johnson Tests of Cognitive Abilities and Tests of Achievement () during in-person testing in a clinical research office conducted by research staff under the supervision of the study neuropsychologist. Tests were administered according to the standardized guidelines and data were entered into the publisher’s scoring program to generate composite scores used in this analysis. Brief Intellectual Ability (BIA) is an age-normalized composite score derived from the Oral Vocabulary, Number Series, and Verbal Attention subtests. Brief Achievement (ACHBRF) is an age-normalized composite score derived from the Letter-Word Identification, Applied Problems, and Spelling subtests. Published reliability for the BIA and ACHBRF are.92 to.95 and.96 to.97, respectively, across our analysis age range ().

2.3.4 Cardiorespiratory fitness assessment (VO2max)

VO2max was estimated using the Parvo Medics True One 2400 Metabolic Measurement System () which controlled a recumbent cycle ergometer in a dedicated physiological assessment laboratory at NKI. Participants exercised at a linearly increasing workload while their heart rate, exhaled CO2, and O2 were analyzed. The assessment was terminated when users met ≥ 90% of their age-related heart rate maximum (220-age) and a respiratory exchange ratio (CO2:O2 ratio; RER) ≥ 1.02, or voluntarily terminated the session.

2.4 Data analyses

2.4.1 cpCST metrics

2.4.1.1 Preprocessing

Raw stimulus coordinate data were preprocessed to correct for deviations caused by a crash during the continuous phase (n = 13 crashes). Crashes, identified as stimulus eccentricity exceeding ± 80% of the distance from the center to the edge of the screen were removed, and the removed data were reconstructed using piecewise polynomial interpolation (PCHIP) to ensure smooth continuity. Participants with two or more crashes were classified as outliers and removed (n = 5).

2.4.1.2 Instantaneous reaction time (iRT)

To quantify temporal responsiveness during task performance, we computed an instantaneous reaction time (iRT) measure using dynamic time warping (DTW). This approach captured continuous time-varying latencies between stimulus and response movements by analyzing the x-coordinate (time) position vectors of both the stimulus object and user positions. The DTW algorithm identified the optimal alignment between these time series, producing a warp path representing temporal correspondence (See Figure 4). By multiplying x-coordinate distances by the sampling rate, we derived latency estimates for each timepoint, providing a highly granular measure of response latency. iRT computations were performed using custom Julia scripts and the DynamicAxisWarping.jl package ().

FIGURE 4

We then computed the mean iRT for each participant, and forwarded these to subsequent analyses.

2.4.2 Flanker metrics

Accuracy and reaction time was recorded for each trial. Participants with accuracy below 80% across all trial conditions were classified as outliers and removed (n = 15). For each participant, incorrect responses were removed from further analysis. For correct trials, anticipatory RTs, defined as RTs faster than 200 ms, as well as RTs more than 2.5 SD longer than the participant’s mean were also removed from further analysis.

We computed the following metrics: mean reaction time for congruent (conRT) and incongruent (incRT) trials, and the standard flanker congruency effect (I-C; incongruent RT - congruent RT). These values were then forwarded for additional analysis.

2.4.3 Reliability in cpCST and flanker

We computed split-half reliability estimates for both the cpCST iRT and the flanker task response times (conRT, incRT, and I-C). To estimate split-half reliability and generate population-level confidence intervals, we used a bootstrap procedure (). In each bootstrap iteration, participants were sampled with replacement, and split-half reliability was computed using the permutation procedure, below.

Split-half reliability in each iteration, trials were randomly permuted and split into two halves. The aggregated mean was computed for each half, and the Pearson correlation between half-scores was calculated. The Spearman–Brown prophecy formula (; ) was applied to correct the correlation, providing an estimate of full-test reliability. This process was repeated 1,000 times, and the average split-half reliability was reported. The split-half approach provides an index of internal consistency by estimating how well two randomly chosen halves of the test relate to each other, scaled to reflect full-test reliability.

The resulting distribution of bootstrap estimates was used to derive 95% confidence intervals (2.5th and 97.5th percentiles).

All reliability estimates were computed using custom Python code, with bootstrap iterations parallelized using Joblib for computational efficiency. Random seeds were fixed to ensure reproducibility.

2.4.4 Temporal efficiency in cpCST and flanker

2.4.4.1 Stability curves

To evaluate the temporal efficiency of each task metric, we assessed how well early portions of the task captured participants’ overall response time (RT) profiles. For each participant, we computed the mean RT separately for each task and condition using only the first n minutes of task data (e.g., first 1, 2, 3, … 9 min). We then correlated these truncated means with the corresponding means computed using the full duration of the corresponding task. This yielded a curve of similarity (Pearson’s r) as a function of data collection time, providing an estimate of how quickly stable RT estimates emerge for cpCST iRT and flanker-based RT metrics.

2.4.4.2 Comparison of stability curves

Statistical comparison between task stability curves for cpCST and flanker trial types was performed using Steiger’s Z-test for dependent correlations with one variable in common (). For each time point (1, 2, 3,. 9 min), we compared the correlation between the truncated and full dataset for the cpCST iRT against the corresponding correlation for each flanker task condition. This approach appropriately accounts for the repeated measures nature of the comparison, estimating the covariance between correlations and compensating for the correlation between the truncated measures (cpCST and flanker). This provides a more conservative and accurate assessment than treating the correlations as independent (). A significant Z-statistic indicates that one task achieves temporal stability more efficiently than the other at that specific time point.

2.4.5 Predictive validity

To evaluate the predictive validity of the cpCST’s instantaneous reaction time (iRT), we conducted a series of regression analyses. Specifically, we examined whether the participants’ iRT could predict performance on proximal experimental measures of inhibitory control and attention (flanker task outcomes), distal clinical measures of cognitive performance (Woodcock-Johnson Cognition and Achievement composite scores), and a measure of central nervous system health and plasticity (VO2max). For each outcome variable, separate regression models were fitted using the mean iRT from the cpCST. We further explored the role of age, repeating these regression analyses both with and without age as a covariate in the models.

3 Results

Participants (N = 166) ranged in age from 18 to 76 years (M = 51.61, SD = 16.36) and reported 12 to 20 years of formal education (M = 15.81, SD = 2.11). The sample was 66% female (n = 110) and 34% male (n = 56). In terms of race, 81% identified as White (n = 134), 10% as Black or African American (n = 16), 5% as Asian (n = 8), 2% as American Indian or Alaska Native (n = 3), and 3% as multiracial (n = 5). Regarding ethnicity, 86% were Not Hispanic or Latino (n = 143), 13% were Hispanic or Latino (n = 22), and 0.6% preferred not to answer (n = 1).

3.1 Reliability of cpCST iRT and flanker outcomes

For the cpCST, the bootstrap-based estimate of population split-half reliability was high [r = 0.9993; 95% CI (0.999, 1.0)]. Split-half reliabilities were also strong for flanker conRT [0.9846; 95% CI: (0.9824–0.9868)] and incRT [r = 0.9752; 95% CI: (0.9725–0.9780)]. Although the split-half reliability for cpCST iRT was statistically greater than the flanker conRT and incRT (p < 0.05), the absolute difference (e.g., 0.9993 vs. 0.9842) is not likely meaningful.

We also assessed the reliability of the standard flanker congruency effect (I-C). The bootstrap-based reliability estimate for this difference score was significantly lower than the cpCST iRT or flanker conRT and incRTs [r = 0.8596; 95% CI: (0.8389–0.8805)].

3.2 Age and sex differences in cpCST and flanker measures

To examine potential individual differences in the primary outcome measures, we performed a series of multiple regression analyses examining the impact of age on the cpCST and flanker measures. See Figure 5 for scatterplots of flanker RTs, cpCST motor stability threshold (MST) and instantaneous reaction time (iRT) as a function of age.

FIGURE 5

3.2.1 cpCST measures

Age significantly predicted the Motor Stability Threshold [MST; B = −0.0025, p < 0.001; F(3, 142) = 57.17, p < 0.001, R2 = 0.55]. For instantaneous reaction time (iRT), age was not a significant predictor [B = 0.0006, p = 0.257; F(2, 143) = 3.36, p = 0.038, R2 = 0.05]. These results indicate that the calibration procedure effectively adjusted for well-documented age-related slowing throughout adulthood.

3.2.2 Flanker measures

Age significantly predicted conRT [B = 2.35, p < 0.001; F(2, 143) = 19.99, p < 0.001, R2 = 0.22] and incRT [B = 2.43, p < 0.001; F(2, 143) = 11.94, p < 0.001, R2 = 0.14]. However, age was not a significant predictor of the I-C congruency effect [B = 0.08, p = 0.755; F(2, 143) = 0.07, p = 0.933, R2 = 0.001].

3.3 cpCST iRT predictive validity

To examine the relationship between cpCST iRT and each of our predicted metrics (Flanker, WJ, and VO2max), we conducted a series of linear regression analyses, both with and without age as a covariate.

3.3.1 Flanker features

Continuous Performance Critical Stability Task iRT significantly predicted conRT [B = 112.48, p = 0.049; F(2, 143) = 20.96, p < 0.001, R2 = 0.23] and incRT [B = 171.91, p = 0.024; F(2, 143) = 14.02, p < 0.001, R2 = 0.16]. However, iRT did not significantly predict the I-C congruency effect [B = 59.42, p = 0.112; F(2, 143) = 1.33, p = 0.268, R2 = 0.02]. When age was included in the models, iRT continued to significantly predict conRT and incRT, while still failing to predict the I-C congruency effect.

Combined, these findings suggest that the cpCST iRT is more closely associated with the response generation aspects of flanker task performance rather than the inhibition of conflicting responses.

3.3.2 WJ brief intellectual ability and WJ brief achievement

Instantaneous reaction time significantly predicted WJ Brief Intellectual Ability [BIA; B = −1.79, p = 0.004; F(2, 143) = 8.98, p < 0.001, R2 = 0.11] and WJ Brief Achievement [ACH; B = −1.42, p = 0.011; F(2, 134) = 8.74, p < 0.001, R2 = 0.12]. Faster iRT was associated with higher ability and achievement scores. Including age in the models did not eliminate these associations, suggesting that the relationships between iRT and the WJ outcome measures were not driven by age.

3.3.3 VO2max

Mean iRT significantly predicted VO2max [B = −9.94, p = 0.010; F(2, 143) = 34.51, p < 0.001, R2 = 0.33]. When controlling for age, iRT remained a significant predictor of VO2max, demonstrating an association of faster reaction time speed with better aerobic capacity, beyond age-related effects.

3.4 Temporal efficiency

The statistical comparison of task stability curves described how well early segments of the task captured participants’ full-task response time (RT) characterizations. The correlation for each mean cumulative (1–9) minute segment of each task’s RT features are plotted below in Figure 6. Even 1 min of iRT data shows very good correlation with the full 10 min assessment (r = 0.87), and by the second minute the correlation with the full sample reached r = 0.94. The flanker Congruent and Incongruent RTs also performed well, though somewhat less well than the iRT. The I-C congruency contrast performed less well than either the iRT or the base flanker features. Locations denoted by a dot on each line show where the correlations for the flanker-based RT features are significantly lower than iRT, using Steiger’s Z-test for dependent correlations with one variable in common ().

FIGURE 6

4 Discussion

We introduced the Continuous Performance Critical Stability Task, offering high temporal precision of continuous psychomotor control across the lifespan. This report provides preliminary evidence for the reliability, predictive validity, and temporally efficiency of the cpCST – a potentially valuable complement to existing attention assessment paradigms. Below, we summarize key methodological innovations and psychometric properties, followed by implications for future research and clinical applications.

4.1 Methodological innovations

The cpCST incorporates three central methodological innovations.

High-density behavioral sampling (30 Hz) captures behavior at a granularity not possible with traditional discrete-response continuous performance tasks, which as noted above, typically sample at rates of 0.1–1 Hz [every 1–10 s; cf (; ; ; )]. The enhanced temporal resolution provides data ideally suited to integrate with other data modalities such as EEG and physiological metrics - allowing sophisticated analyses of attentional stability and variability.

We also created a novel instantaneous reaction time measure, which estimates the temporal lag between the movement of a central stimulus object and the participant’s response to adjust to that movement. This approach estimates response time with high precision, reliability, and excellent temporal efficiency.

Additionally, the cpCST utilizes a hybrid design that combines adaptive calibration and subsequent fixed-difficulty assessments. Integrating the strengths of adaptive and fixed-difficulty paradigms provides individualized task difficulty while avoiding issues common in fully adaptive methods, such as oscillatory artifacts or instability (; ). It may also obviate the need for alternative task forms across groups with disparate baseline functioning, or in highly heterogeneous samples such as in aging, developmental, or lifespan studies.

4.2 Psychometric properties

The cpCST yielded high reliability estimates, with bootstrap-based split-half reliability greater than 0.999. High-density sampling and individualized calibration likely contributed to this reduced measurement error, facilitating the rapid detection of subtle individual differences (r > 0.9 after 1 min of data). This reliability may be especially advantageous in longitudinal studies or in interventions examining modest performance changes.

Age invariance is a notable strength of the cpCST. Although motor stability thresholds (MST) and traditional reaction time measures from the flanker task exhibited expected age-related slowing, cpCST’s iRT was stable across age. By calibrating task difficulty to each individual’s sensorimotor capacity, the cpCST appeared to effectively isolate attentional control from baseline sensorimotor function. This makes the task especially suitable in lifespan cognitive assessments, circumventing the need for distinct age-specific task versions.

The cpCST also exhibited robust validity across multiple domains. Significant associations with flanker conRT and incRT suggest convergent validity with aspects of attentional control. However, the lack of association with the flanker congruency effect may indicate that the cpCST primarily captures tonic aspects of attention (e.g., vigilance, sustained focus) rather than the application of inhibitory control processes. Head-to-head comparisons of cpCST performance metrics with established measures of vigiliance, sustained attention, and other dimensions of attentional control, while outside the scope of this analysis, are nevertheless warranted to characterize cpCST construct validity.

The temporal efficiency of the cpCST was also notable. Over 95% of participants reached asymptotic performance within the first 1.5 min of the calibration phase. Within 2 min of the continuous phase, the cpCST iRT exceeded an r = 0.9 correlation with full task performance. By comparison, the flanker trial types needed roughly 5 min of data to reach this level of association with the full flanker sample. This suggests the potential for cpCST to reduce task administration time without significant loss of information.

4.3 Implications and future directions

We identified significant predictive relationships across a broad range of domains, encompassing individual differences in low-level physiological functioning (VO2max), reaction time in a traditional cognitive task (flanker), and even global estimates of intellectual ability and achievement (WJ Brief Intellectual Ability, Brief Achievement). While speculative, this remarkable range of associations suggests that the cpCST may tap one or more central aspects of neurocognitive functioning. Future research to better contextualize the cpCST amongst the existing constellation of cognitive assessments will likely be of high value.

The central features of the cpCST position it as a promising tool for research and clinical settings. Its high temporal resolution enables tighter integration with physiological measures (e.g., EEG, fMRI, heart rate, skin conductance), facilitating exploration of neural mechanisms underlying moment-to-moment attention variability, and “brain-body” interactions. Additionally, its individualized calibration method is likely to prove valuable in heterogeneous clinical populations or lifespan studies, as it reduces confounds related to sensorimotor speed differences or ceiling/floor effects.

The task’s temporal efficiency and straightforward administration suggest suitability for large-scale assessments, longitudinal monitoring, and remote or mobile implementations. Future studies should explicitly evaluate cpCST’s sensitivity to attentional changes resulting from interventions (e.g., sleep deprivation, stimulant medication, cognitive training) and establish its utility in diverse clinical populations (e.g., ADHD, TBI, MCI). Additionally, as illustrated in Figure 7, the high density sampling may allow detection of subtle behavioral dynamics not captured with discrete response paradigms - which may not only contribute to the cpCST’s relatively high temporal efficiency and reliability, but also allow for new insights into attentional dynamics.

FIGURE 7

Finally, the rich, high-density behavioral data generated by the cpCST is well-suited for computational modeling approaches, such as drift diffusion models or Bayesian frameworks. Future work could leverage these modeling techniques to better characterize the attentional process dynamics captured by the cpCST.

4.4 Limitations

Several limitations should be acknowledged. Although predictive validity and split-half reliability were established, the cpCST’s sensitivity to intervention-induced attentional changes remains to be validated. As noted in Psychometric Properties, above, cpCST task performance was not directly compared to a full range of established measures of sustained attention or attentional control. Future work that comprehensively reviews the theoretical positioning of widely adopted and emerging attention tasks and provides psychometric evaluation via head-to-head empirical evidence for both shared and unique behavioral features would provide useful information to guide research advances in theoretical and practical applications. Our current analysis age range (18–76 years) is substantial, but was undertaken as a preliminary convenience sample; larger samples that include evaluation of efficacy and validity in younger and older individuals require further examination. Likewise, this is a community-based normative sample and psychometric properties should be evaluated across different clinical populations. Given the cross-sectional nature of our sample, we can only establish internal reliability through bootstrap methods. Future work is needed to examine test-retest reliability under frameworks like the intraclass correlation coefficient [ICC; (; )].

Additionally, while we include summary evidence of calibration feasibility and sensitivity, a full psychometric evaluation of the calibration phase (e.g., MST distributions, convergence dynamics, and predictive validity) is beyond the scope of this initial paper and will be presented in a companion manuscript.

Finally, while high-density behavioral sampling offers analytical richness, the relative complexity of calculating iRT using dynamic time warping (DTW) may present obstacles to widespread adoption. To address this, we will provide streamlined and containerized analysis pipelines on GitHub. Developing accessible pipelines and normative databases will be essential for broader clinical adoption and research utilization.

5 Conclusion

The Continuous Performance Critical Stability Task introduces methodological advances in the assessment of attention. Its exceptional reliability, age invariance, predictive validity, and temporal efficiency address limitations in existing measures. Future validation efforts integrating physiological measures, computational modeling, and diverse clinical applications will further establish cpCST’s utility as an essential tool for attention research and assessment.

Statements

Data availability statement

The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found below: the datasets generated for this study can be found in the NKI-Rockland Sample data resource https://rocklandsample.org/. Access to raw data requires completion of a data use agreement. Analysis datasets used for the current study are available from the corresponding author upon reasonable request.

Ethics statement

The studies involving humans were approved by Nathan S. Kline Institute for Psychiatric Research - New York State Office of Mental Health- Institutional Review Board. The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.

Author contributions

AM-B: Conceptualization, Formal analysis, Methodology, Project administration, Supervision, Visualization, Writing – original draft, Writing – review & editing. DG-B: Data curation, Formal analysis, Investigation, Visualization, Writing – original draft, Writing – review & editing. KG: Data curation, Investigation, Visualization, Writing – original draft, Writing – review & editing. OR: Data curation, Investigation, Visualization, Writing – review & editing. EG: Data curation, Formal analysis, Validation, Visualization, Writing – review & editing. MM: Conceptualization, Funding acquisition, Methodology, Supervision, Writing – review & editing. SC: Conceptualization, Formal analysis, Funding acquisition, Methodology, Project administration, Supervision, Visualization, Writing – original draft, Writing – review & editing.

Funding

The author(s) declare that financial support was received for the research and/or publication of this article. This work was supported by the National Institute of Mental Health (R01MH124045; The NKI Rockland Sample II: An Open Resource of Multimodal Brain, Physiology & Behavior Data from a Community Lifespan Sample). Additional support was provided by the New York State Office of Mental Health and the Nathan Kline Institute institutional core services. Author salaries were supported by New York State (AM-B, SC, MM, EG) and by the NIH grant R01MH124045 (DG-B, OR, KG).

Acknowledgments

We acknowledge the important support of the entire Nathan Kline Institute – Rockland Sample team of investigators, research support staff, NKI Scholars and, most importantly, the community member participants who volunteer time and energy to advance scientific knowledge and benefit others.

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The authors declare that no Generative AI was used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

Summary

Keywords

sustained attention, sensorimotor integration, reaction time, behavioral assessment, lifespan development, adaptive testing, task reliability, individual differences

Citation

MacKay-Brandt A, Garcia-Barnett D, Gan KX, Ripley O, Gazes E, Milham M and Colcombe S (2025) cpCST: a new continuous performance test for high-precision assessment of attention across the lifespan. Front. Psychol. 16:1640417. doi: 10.3389/fpsyg.2025.1640417

Received

03 June 2025

Accepted

22 August 2025

Published

23 September 2025

Volume

16 - 2025

Edited by

Richard A. Abrams, Washington University in St. Louis, United States

Reviewed by

Giulio Contemori, University of Padua, Italy

Sofia Abrevaya, National Scientific and Technical Research Council (CONICET), Argentina

Updates

Copyright

*Correspondence: Stan Colcombe,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics