METHODS article

Front. Appl. Math. Stat., 25 January 2021

Sec. Mathematics of Computation and Data Science

Volume 6 - 2020 | https://doi.org/10.3389/fams.2020.611805

Causal Analysis of Health Interventions and Environments for Influencing the Spread of COVID-19 in the United States of America

  • 1. School of Public Health, The University of Texas Health Science Center at Houston, Houston, TX, United States

  • 2. Tulane Center of Biomedical Informatics and Genomics, Deming Department of Medicine, School of Medicine, Tulane University, New Orleans, LA, United States

Abstract

Given the lack of potential vaccines and effective medications, non-pharmaceutical interventions are the major option to curtail the spread of COVID-19. An accurate estimate of the potential impact of different non-pharmaceutical measures on containing, and identify risk factors influencing the spread of COVID-19 is crucial for planning the most effective interventions to curb the spread of COVID-19 and to reduce the deaths. Additive model-based bivariate causal discovery for scalar factors and multivariate Granger causality tests for time series factors are applied to the surveillance data of lab-confirmed Covid-19 cases in the US, University of Maryland Data (UMD) data, and Google mobility data from March 5, 2020 to August 25, 2020 in order to evaluate the contributions of social-biological factors, economics, the Google mobility indexes, and the rate of the virus test to the number of the new cases and number of deaths from COVID-19. We found that active cases/1,000 people, workplaces, tests done/1,000 people, imported COVID-19 cases, unemployment rate and unemployment claims/1,000 people, mobility trends for places of residence (residential), retail and test capacity were the popular significant risk factor for the new cases of COVID-19, and that active cases/1,000 people, workplaces, residential, unemployment rate, imported COVID cases, unemployment claims/1,000 people, transit stations, mobility trends (transit), tests done/1,000 people, grocery, testing capacity, retail, percentage of change in consumption, percentage of working from home were the popular significant risk factor for the deaths of COVID-19. We observed that no metrics showed significant evidence in mitigating the COVID-19 epidemic in FL and only a few metrics showed evidence in reducing the number of new cases of COVID-19 in AZ, NY and TX. Our results showed that the majority of non-pharmaceutical interventions had a large effect on slowing the transmission and reducing deaths, and that health interventions were still needed to contain COVID-19.

Introduction

As of August 25, 2020, the number of cumulative cases of COVID-19 in the US exceeded 5,727,107 and included 170,305, deaths (John Hopkins Coronavirus Resource Center, https://coronavirus.jhu.edu/MAP.HTML), thus causing a devastating public health and economic crisis. Since the number of new cases in the US remains high (36,339 in the US on August 25, 2020) ((John Hopkins Coronavirus Resource Center, https://coronavirus.jhu.edu/MAP.HTML), curbing the spread of COVID-19 is urgently needed []. There is increasing recognition that many geographic, economic and environmental factors contribute to the outbreak of COVID-19. In the absence of vaccines and specifically effective medications, non-pharmaceutical public health interventions and personal hygiene practices are the only options to slow the spread of COVID-19 [, ]. The effects of the different factors and intervention measures on the spread of COVID-19 vary. Identifying key factors that most contribute to the rapid spread of COVID-19, and accurately estimating the potential impact of different non-pharmaceutical measures for containing COVID-19 are crucial for planning the most effective interventions to curb the spread of COVID-19 [].

The widely used statistical methods for COVID-19 epidemiological factor analysis and evaluation of intervention measures include correlation analysis [, ], regression [, ], logistic regression [] and a transmission dynamic model coupled with a linear model []. The most examined scalar factors consist of underlying health conditions such as high blood pressure, diabetes, stroke, cardiac or kidney diseases, and aging individuals [, , ], atmospheric temperature [], age, gender, ethnicity, and population density [, ], airflow [], and socioeconomics such as median income [, ].

The most explored non-pharmaceutical public health interventions and digital technologies for curbing the spread of COVID-19 include social distancing, case isolation and quarantine as well as closuring borders, schools travel restrictions, use of face-masks, and testing [] and population surveillance, case identification, contact tracing, mobility data collection, and communication technology, which utilize billions of mobile phones and large online datasets to provide information for the evaluation of intervention strategies and to strengthen the curb of the spread of COVID-19 [].

Although association analysis is of great importance for curbing the spread of COVID-19, association measures dependence between two variables or two sets of variables in the data, and use the dependence for prediction and evaluation of the effects of environmental, social-economic factors and public health interventions on the spread of COVID-19 [, ]. It is well recognized that association analysis is not a direct method to discover the causal mechanism of complex diseases. Association analysis may detect superficial patterns between intervention measures and transmission variables of COVID-19. Its signals provide limited information on the causal mechanism of the transmission dynamics of COVID-19 []. Association analysis has been a major paradigm for statistical evaluation of the effects of influencing factors and health interventions on the spread of COVID-19 []. Understanding the transmission mechanism of COVID-19 based on association analysis remains elusive. The question to uncover the transmission mechanisms of COVID-19 is causal in nature.

Distinguishing causation from association is an age-old problem. Methods for causation analysis that is one of the most challenging problems in science and technology need to be developed as an alternative to association analysis []. A number of researchers have performed causal analysis of COVID-19 to evaluate the causal effects of mobility, awareness, and temperature [], social distancing [], mobility [], herd immunity [], and mask use []. However, most causal analysis of COVID-19 have treated time series data as pseudo-cross-sectional data. In some cases, causal analysis of COVID-19 treated the data as time series; time series was assumed stationary. In practice, the number of new cases and the number of deaths from COVID-19 were nonstationary time series in most cases []. The environmental, social-economic and geographic factors, and intervention measures include two types of data: scalar variables and time series (stationary or nonstationary) variables.

The purpose of this paper is to develop a general framework for the causal analysis of COVID-19 in the US. The number of new cases and deaths from COVID-19 are taken as response variables. The factors and intervention measures are taken as potential causal variables. If the factor and intervention variables are scalar variables, the additive noise models (ANMs) [] are used to test for causation between the response variable and potential causal variable where the number of new cases or deaths should be averaged over time. Most intervention measures are time series data. An essential difference between time series and cross-sectional data is that the time series data have temporal order, but cross-sectional data do not have any order. As a consequence, the causal inference methods for cross sectional data cannot be directly applied to time series data. Basic tools in statistical analysis are the raw of large numbers and the central limit theorem. Applications of these tools usually assume that all moment functions are constant. When the moment functions of the time series vary over time, the raw of large numbers and the central limit theorem cannot be applied. In order to use basic probabilistic and statistical theories, the nonstationary time series must be transformed to stationary time series [].

A widely used concept of causality for time series data is Granger causality [

,

]. Underlying the Granger causality is the following two principles:

  • (1) Effect does not precede the cause in time;

  • (2) The effect series contains unique causal series information which is not present elsewhere.

The multivariate linear Granger causality test will be used to test causality between the number of new cases and deaths from COVID-19 and environmental, economic and intervention time series variables []. The proposed ANMs and multivariate linear Granger causality analysis methods are applied to the surveillance data of lab-confirmed COVID-19 cases in the US, UMD data, and Google mobility data from March 5, 2020 to August 25, 2020 in order to evaluate the contributions of social-biological factors, economics, the Google mobility indexes, and the rate of virus testing to the number of the new cases and number of deaths from COVID-19. Data and software for implementing the algorithms for causal analysis can be downloaded from our website https://sph.uth.edu/research/centers/hgc/software/xiong/.

Material and Methods

Nonlinear Additive Noise Models for Bivariate Causal Discovery

The ANMs are used for identifying causal effect of a factor or an intervention measure on the number of new cases or an intervention measure [, ]. Assume no confounding, no selection bias, and no feedback. Let Y be the average number of new cases or new deaths from COVID-19 and X be a scalar factor or an intervention measure such as gender, population density, ethnic group, among others. Consider a bivariate additive noise model where Y is a nonlinear function of X and independent additive noise :where X and are independent. Then, the density is said to be induced by the additive noise models (ANM) from X to Y []. In some cases, we may have the following alternative direction of the ANMs: :where Y and are independent. If the density is induced by the ANM , but not by the ANM , then the ANM is identifiable.

Assume that n + m state data were sampled. Divide the dataset into a training data set by specifying for fitting the model and a test data set for testing the independence, where n is not necessarily equal to m.

Procedures for using the ANM to assess causal relationships between two variables are summarized below [

].

  • Step 1. Regress Y on X using the training dataset and non-parametric regression methods:

  • Step 2. Calculate the residual using the test dataset and test whether the residual is independent of causal X to assess the ANM .

  • Step 3. Repeat the procedure to assess the ANM .

  • Step 4. If the ANM in one direction is accepted and the ANM in the other is rejected, then the former is inferred as the causal direction.

There are many non-parametric methods that can be used to regress Y on X or regress X on Y. For example, we can use neural networks [], smoothing spline regression methods [], B-spline [] and local polynomial regression []. In this paper, the smoothing spline regression method was used to fit the regression models.

Covariance can be used to measure association but cannot be used to test independence between two variables with a non-Gaussian distribution (https://en.wikipedia.org/wiki/Correlation_and_dependence). A covariance operator that is a generalization of the finite dimensional covariance matrix to infinite dimensional feature space can be used to test for independence between two variables with arbitrary distributions. Specifically, we will use the Hilbert-Schmidt norm of the cross-covariance operator or its approximation, the Hilbert-Schmidt independence criterion (HSIC) to measure the degree of dependence between the residuals and potential causal variable and test for their independence [, ].

The covariance operator can be defined aswhere are any nonlinear functions and is the covariance operator and is an inner product in the Hilbert space. The covariance operator is positive symmetric, which implies linearity and continuity. In addition, the covariance operator maps spaces in their dual spaces []. The Hilbert-Schmidt norm of the covariance operator can be used as criterion for assessing independence between two random variables and is called the Hilbert-Schmidt independence criterion (HSIC). The Hilbert-Schmidt norm of the centered covariance operator is defined aswhere is the Hilbert-Schmidt norm.

We know that (Wang et al., 2018). if and only if and are independent. can be approximated by.where is a sample size, and are dimensional kernel matrices and . We used the Gaussian kernel: and polynomial kernel. The order- polynomial kernel is defined as . The fourth and sixth order polynomial kernels were used in this analysis. To test independence between the potential cause and residual , we calculated as follows.

In summary, the general procedure for testing independence between the average number of new cases or new deaths and the scalar factor or intervention measure is given as follows [

,

]:

  • Step 1: Divide a data set into a training data set for fitting the model and a test data set for testing the independence.

  • Step 2: Use the training data set and any non-parametric regression methods

    • (a)Regress on : ,

    • (b)Regress on : .

  • Step 3: Use the test data set and any non-parametric regression methods that fits the test data set to predict residuals:

    • (a)

    • (b)

  • Step 4: Calculate the dependence measures and .

  • Step 5: Infer causal direction:

If

then causal direction is undecided.

We do not have closed analytical forms for the asymptotic null distribution of the HSIC and hence it is difficult to calculate the p-values of the independence tests. To solve this problem, the permutation/bootstrap approach can be used to calculate the p-values of the causal test statistics. The null hypothesis is no causations and (Both and are dependent, and and are dependent).

Calculate the test statistic

Assume that the total number of permutations is . For each permutation, we fix and randomly permutate Then, fit the ANMs and calculate the residuals and test statistic . Repeat above procedures times. The p-values are defined as the proportions of the statistic (computed on the permuted data) greater than or equal to (computed on the original test data ).

Each state was a sample. Since the sample sizes were small (only 50), the p-value for declaring significance was 0.05 without Bonferroni correction for multiple comparison.

Multivariate Linear Granger Causality Test

Before performing multivariate linear Grander causality test, we first need to transform nonstationary time series to stationary time series.

Consider an -variable VAR with lags:where is a dimensional vector, is a mean, the are coefficient matrices and dimensional residual vector is assumed to have mean zero with no autocorrelation but can be correlated across equations .

Vector error correction model (VECM) consists of first differences of cointegrated variables, their lags, and error correction terms:where matrixes and are functions of matrices .

When two non-stationary variables are cointegrated, the VAR model should be augmented with an error correction term for testing the Granger causality (Engle and Granger, 1987).

The VECM can be reduced towhere

Consider two non-stationary time series, and . Let .

Suppose that and are cointegrated with the residuals . The VECM model for testing the Granger causality is given bywhere and are and dimensional vectors of intercept terms, respectively, and are and dimensional matrices of lag polynomials, respectively, and are and dimensional coefficient vectors for the error correction term respectively. The lag length was selected using the two-stage procedure [].

There are four different cases of causal relationships between two vectors of time series

and

[

].

  • (1) If is significantly different from the zero, while shows no significantly different from zero, then there exists a unidirectional Ganger causality from time series to ;

  • (2) If is significantly different from zero, while shows no significantly difference from zero, then there exists a unidirectional Ganger causality from to ;

  • (3) If both coefficients and are significantly different from zero, then there exists bidirectional Granger causality between and ;

  • (4) If both coefficients and are not significantly different from zero, then and are not rejected to be independent.

The four statements imply that Ganger causal relationships between

and

depend on the coefficients

and

Therefore, the null hypotheses for testing the Ganger causality between

and

are.

  • (1)

  • (2) and

  • (3) Both and and

Likelihood ratio tests for multivariate Granger causality are given by.

  • (1) The likelihood ration statistics for testing the null hypothesis: is

which is asymptotically distributed as a central

under the hull hypothesis

.

  • (2)The likelihood ration statistics for testing the null hypothesis: is

which is asymptotically distributed as a central

under the hull hypothesis

.

  • (3)The likelihood ration statistics for testing the null hypothesis: and and is

which is asymptotically distributed as a central

under the hull hypothesis

and

.

The total number of variables to be tested was 18. The p-value for declaring significance after Bonferroni correction was 0.0028.

Data Collection

Data on the number of new cases and new deaths of COVID-19 across the 50 states in the US were obtained from John Hopkins Coronavirus Resource Center (https://coronavirus.jhu.edu/MAP.HTML). Google mobility indexes were downloaded from Google COVID-19 Community Mobility Reports (https://www.google.com/covid19/mobility/). Comprehensive data and insights on COVID-19’s impact on mobility, economy, and society were downloaded from the University of Maryland COVID-19 Impact Analysis Platform (https://data.covid.umd.edu) [, ]. All data were collected from March 5, 2020 to August 25, 2020.

Results

Test for Scalar Potential Causes

The scalar variables tested for causation of the new cases and deaths from COVID-19 in the US included the number of contact tracing workers per 100,000 people, percent of population above 60 years of age, median income, population density, percentage of African Americans, percentage of Hispanic Americans, percentage of males, employment density, number of points of interests for crowd gathering per 1,000 people, number of staffed hospital beds per 1,000 people, and number of ICU beds per 1,000 people. The number of new cases and deaths were averaged over time. Each state was a sample. Since the sample sizes were small, the p-value for declaring significance was 0.05 without Bonferroni correction for multiple comparison. The p-values for testing 11 scalar potential causes of the number of new cases and deaths from COVID-19 in the US using both Gaussian kernel and polynomial kernels were summarized in Table 1 where minimum of three p-values using Gaussian kernel and fourth order and sixth order polynomial kernels were listed and only one causal direction was observed. We observed from Table 1 that population density (minimum of p-value < 0.0002, which was due to Gaussian kernel), percentage of males (minimum of p-value < 0.03, which was due to Gaussian kernel) and Percentage of Hispanic Americans (minimum of p-value < 0.0325, which was due to sixth order polynomial kernel) showed significant evidence of causing the spread of COVID-19. Employment density (minimum of p-values < 0.0223, which was due to sixth polynomial kernel), Percentage of African American (minimum of p-values < 0.024, which is due to Gaussian kernel), population density (minimum of p-values < 0.025, which was due to Gaussian kernel) and Percentage of males (minimum of p-values <0.0377, which was due to fourth order polynomial) showed significant evidence of causing deaths due to COVID-19. percentage of Hispanic Americans (minimum of p-value < 0.052, which was due to sixth order polynomial kernel) were close to significance level 0.05 for causing death.

TABLE 1

Risk factorp-value
Gaussian kernelPolynomial (order 6)Polynomial (order 4)Minimum
New casesDeathsNew casesDeathsNew casesDeathsNew casesDeaths
Percent of population above the age of 60 years0.46340.15300.34720.11300.25320.15470.25320.1130
Median income0.11090.07600.12020.06540.17130.18320.11090.0654
Percentage of African Americans0.65260.02400.93260.07100.34590.03430.34590.0240
Percentage of Hispanic Americans0.05750.06400.03250.05200.04220.14730.03250.0520
Percentage of males0.03000.14400.04030.05420.04010.03770.03000.0377
Population density0.00020.02500.03270.03590.02650.09520.00020.0250
Employment density0.45710.05900.35360.02230.34850.02480.34850.0223
Number of staffed hospital beds per 1,000 people0.67320.31300.53520.32240.65060.23740.53520.2374
Number of ICU beds per 1,000 people0.51340.48600.75160.44380.61140.90370.51340.4438
Number of contact tracing workers per 100,000 people0.42030.81900.41020.72150.39070.73360.39070.7215

p-values for testing 10 scalar potential causes of the number of new cases and deaths of COVID-19 in the US.

Population density was an important risk factor for both the spread and death from COVID-19. High density resulted in closer contact, stronger interaction among residents and lower social distancing, which facilitated the spread and increased the death rate from COVID-19 []. However, our results were contradictory with the conclusion of Hamidi et al. []. Some literature also confirmed that high proportion of African Americans caused a high rate of deaths [, 50, 51]. Our results concluded that percentage of Hispanic Americans was a risk factor for the spread and a weak risk factor for death from COVID-19, while the literature showed stronger evidence that Hispanic communities were highly vulnerable to COVID-19 [52].

The second most significant demographic risk factor for the spread of COVID-19 was percentage of males. We found higher COVID-19 morbidity in males than females. However, we did not find higher COVID-19 mortality in males than females.

It was reported that higher COVID-19 mortality in males than females can be due to the following factors [53]. The first factor was higher expression of angiotensin-converting enzyme-2 (ACE two; receptors for coronavirus) in males than females. The second factor was sex-based immunological differences due to sex hormone and the X chromosome.

Test for Granger Causality

Daily mobility and social distancing data from a COVID-19 impacted the analysis platform, including four categories: category A: mobility and social distancing, category B: COVID and health, category C: economic impact, and category D: vulnerable population. A total of 12 temporal metrics in four categories and 12 metrics from the COVID-19 impact analysis platform, six daily Google Community Mobility indexes and protest attendee data that captured real-time trends in movement patterns for each state in the US were included in the analysis to test for Granger causality between these risk factors, health intervention measures and the number of new cases and deaths from COVID-19 across 50 states in the US [, ]. The total number of variables to be tested was 18. The p-value for declaring significance after Bonferroni correction was 0.0028.

All 18 metrics except for protest attendee showed high significance in causing a reduction of the new cases of COVID-19 in 19 less affected states: VT, WY, ME, AK, NH, WV, ND, SD, NM, RI, DE, KY, KS, CT, CO, IA, WA, WI, and MS. Most of these states were less populated. However, although CA was most affected and the most populated state, all 18 metrics except protest attendance showed a strong significance in causing rapid spread of COVID-19 (Table 2 and Supplementary Table S1). To provide complete causal testing information, we listed the values of statistics for testing 18 temporal potential causes of the number of new cases of COVID-19 across 50 states in the US in Supplementary Table S2.

TABLE 2

StateCAFLTXNYGAILAZNJNCTN
Number of cumulative cases673,095605,502586,730430,774258,354224,887199,273190,021157,741145,417
Miles/person7.7E-083.1E-022.3E-039.1E-022.8E-044.9E-047.3E-054.7E-039.2E-065.9E-06
Population1.6E-063.6E-025.1E-036.1E-023.2E-045.2E-051.9E-041.2E-032.5E-042.9E-05
% Change in consumption1.4E-055.5E-024.7E-031.8E-038.7E-041.8E-042.2E-049.6E-071.1E-041.1E-04
Social distancing index8.0E-069.1E-028.8E-034.0E-023.4E-044.6E-045.9E-041.5E-053.5E-032.7E-04
Unemployment claims4.3E-056.8E-021.7E-021.1E-032.3E-038.0E-041.1E-032.2E-074.2E-033.9E-04
Unemployment rate3.7E-082.6E-024.4E-031.7E-014.2E-045.3E-062.4E-042.8E-031.5E-044.8E-05
% Working from home5.2E-073.7E-022.4E-033.4E-022.4E-046.6E-062.2E-052.8E-041.2E-052.0E-06
Active cases/1,000 people2.3E-187.2E-023.1E-065.8E-012.8E-102.3E-057.1E-079.5E-047.5E-143.4E-12
Testing capacity9.0E-057.2E-021.3E-023.0E-022.4E-032.1E-054.1E-041.2E-041.7E-031.8E-04
Tests done/1,000 people1.2E-143.6E-021.1E-051.0E-012.5E-091.5E-037.0E-053.8E-031.3E-073.7E-11
Imported COVID cases1.3E-164.5E-011.1E-042.6E-011.7E-071.3E-035.1E-045.9E-041.7E-064.4E-12
Retail1.3E-059.6E-022.2E-022.9E-024.4E-035.9E-041.4E-033.4E-074.5E-036.1E-04
Grocery1.5E-051.1E-014.1E-032.0E-015.8E-041.5E-052.7E-049.0E-062.6E-044.5E-05
Parks2.4E-088.2E-026.3E-034.1E-023.1E-026.4E-025.6E-034.0E-038.6E-031.7E-05
Transit5.8E-066.4E-027.3E-032.0E-023.3E-042.2E-073.0E-049.2E-061.2E-033.3E-04
Workplaces1.5E-059.6E-032.6E-033.5E-032.1E-055.8E-089.3E-062.2E-051.2E-046.1E-05
Residential7.0E-051.7E-027.2E-039.6E-053.6E-041.7E-076.5E-051.1E-062.8E-041.5E-04
Protest3.6E-019.6E-016.6E-017.7E-018.0E-035.1E-013.8E-011.7E-014.1E-012.3E-03
StateSDNDWVNHHIMTAKMEWYVT
Number of cumulative cases11,55910,2299,3957,1506,9846,6245,6664,3683,6341,572
Miles/person3.7E-081.1E-041.3E-093.8E-149.5E-035.8E-042.8E-058.2E-112.7E-134.8E-14
Population4.2E-085.4E-046.6E-081.7E-182.0E-021.2E-035.4E-041.7E-121.7E-116.8E-16
% Change in consumption2.1E-081.1E-044.6E-103.6E-193.1E-021.4E-047.2E-051.2E-114.2E-107.5E-19
Social distancing index5.7E-059.8E-035.0E-062.8E-193.3E-026.5E-032.4E-031.6E-081.9E-091.6E-19
Unemployment claims5.0E-064.3E-031.8E-063.6E-163.1E-023.3E-031.4E-031.9E-071.0E-082.1E-32
Unemployment rate3.8E-081.3E-032.5E-064.8E-262.9E-022.7E-032.8E-047.1E-188.0E-123.8E-14
% Working from home1.5E-091.3E-051.5E-121.4E-162.3E-031.1E-048.3E-041.3E-112.4E-151.8E-16
Active cases/1,000 people3.9E-093.3E-131.5E-166.2E-271.3E-084.8E-093.2E-076.8E-132.1E-216.9E-15
Testing capacity1.1E-073.4E-055.4E-071.3E-222.1E-023.6E-032.0E-031.7E-146.0E-101.5E-23
Tests done/1,000 people4.5E-074.6E-112.1E-161.9E-123.8E-051.9E-116.3E-097.4E-071.5E-195.7E-13
Imported COVID cases2.7E-076.8E-099.6E-139.0E-327.4E-062.7E-082.0E-061.3E-145.4E-192.5E-13
Retail1.3E-076.1E-034.4E-064.8E-192.9E-022.7E-043.1E-042.0E-095.4E-141.4E-20
Grocery1.3E-043.7E-041.1E-071.7E-222.0E-022.3E-047.5E-044.3E-103.3E-135.6E-20
Parks3.9E-064.9E-041.6E-063.9E-121.9E-021.6E-072.8E-047.8E-071.6E-193.1E-12
Transit3.3E-063.2E-037.4E-091.3E-222.2E-021.6E-045.6E-046.3E-112.0E-133.0E-19
Workplaces3.9E-122.3E-041.2E-071.4E-231.8E-024.6E-044.7E-049.8E-151.9E-092.2E-19
Residential3.5E-091.6E-032.0E-063.0E-241.6E-021.4E-036.3E-044.9E-125.6E-083.1E-22
Protest2.4E-012.6E-019.7E-011.1E-031.3E-022.1E-019.6E-038.0E-028.3E-011.6E-01

p-values testing 18 temporal potential causes of the number of new cases of COVID-19 in the top 10 most affected states and bottom 10 less affected states in the US.

Retail: Retail and recreation, mobility trends for places like restaurants, cafes, shopping centers, theme parks, museums, libraries, and movie theaters.

Grocery: Mobility trends for places like grocery markets, food warehouses, farmers markets, specialty food shops, drug stores, and pharmacies.

Parks: Mobility trends for places like local parks, national parks, public beaches, marinas, dog parks, plazas, and public gardens.

Transit: Transit stations, mobility trends for places like public transport hubs such as subway, bus, and train stations.

Workplaces: Mobility trends for places of work.

Residential: Mobility trends for places of residence.

Test Rate: Ratio of the number of individuals who have taken the virus test over the total population in the region.

Attendee: Number of attendees in the protest.

All 18 metrics showed no significance in causing reduction of the new cases of COVID-19 in Florida. The majority of the 18 metrics did not demonstrate evidence that they can significantly mitigate the spread of COVID-19 in most the affected states such as TX, NY, GA, IL, AZ, NJ, NC, and TN. These 10 states were in the top largest states by population in the US. Public health intervention measures such as closing schools and businesses, avoiding public gatherings, restricting traffic, placing residents to stay-at-home and adherence to guidelines were less well implemented or difficult to implement homogeneously due to large populations and geographical areas [55]. These results also explained why the number of new cases of COVID-19 in these states was high and confirmed by several studies [5658].

Table 3 summarized the ranges of p-values, Supplementary Tables S3, S4 summarized all p-values and values of statistics for testing 18 temporal potential causes of the number of new deaths from COVID-19 across 50 states in the US, respectively. All 18 metrics except for protest attendance showed high significant evidence for causing a reduction of new deaths across 50 states except for Michigan (MI) in the US. Our results suggested that a cascade of causes led to the COVID-19 tragedy in the US.

TABLE 3

Statep-valueStatep-value
Lower boundUpper boundLower boundUpper bound
AK1.47E-183.95E-24MT7.66E-154.41E-25
AL7.45E-111.10E-18NC2.93E-088.95E-18
AR1.17E-121.54E-30ND2.54E-171.04E-20
AZ4.35E-111.29E-27NE5.11E-152.62E-23
CA1.10E-041.54E-13NH6.67E-153.41E-28
CO1.41E-216.44E-26NJ1.80E-031.96E-09
CT6.73E-073.13E-16NM7.20E-071.15E-16
DE1.20E-031.35E-06NV9.00E-091.45E-17
FL2.50E-041.47E-15NY4.30E-021.52E-05
GA6.53E-075.81E-17OH1.82E-082.67E-17
HI1.77E-151.14E-20OK4.29E-071.89E-19
IA6.78E-068.06E-15OR3.37E-154.50E-26
ID1.96E-087.11E-18PA3.86E-151.11E-23
IL3.81E-051.34E-12RI4.26E-092.16E-22
IN1.78E-073.39E-17SC8.97E-086.75E-25
KS8.74E-301.42E-40SD5.04E-186.01E-25
KY1.89E-131.40E-24TN7.28E-092.06E-24
LA7.71E-112.38E-20TX6.22E-072.59E-23
MA1.25E-033.92E-04UT3.90E-138.12E-29
MD3.80E-027.63E-06VA3.39E-172.61E-16
ME2.29E-154.29E-25VT3.39E-071.23E-29
MI0.282.07E-05WA5.67E-163.83E-29
MN2.40E-048.89E-11WI2.21E-104.77E-29
MO1.73E-117.85E-20WV9.61E-157.81E-22
MS2.94E-081.49E-19WY7.93E-262.39E-28

Ranges of p-values for testing 18 temporal potential causes of the number of new deaths from COVID-19 across 50 states in the US.

Table 4 listed the most significant risk factor for the new cases of COVID-19 in each of the 50 states in the US. Active Cases/1,000 People, workplaces, number of tests completed/1,000 people, imported COVD cases, unemployment rate and unemployment claims/1,000 people, mobility trends for places of residence (residential), retail and recreation, mobility trends for places like restaurants, cafes, shopping centers, theme parks, museums, libraries, and movie theaters (retail) and test capacity were the most significant risk factors for the new cases of COVID-19 in 23, 7, 6, 5, 4, 2, 1 and 1 states of the US, respectively.

TABLE 4

StateRisk factorp-valueStateRisk factorp-value
AKTests done/1,000 people6.27E-09MTTests done/1,000 people1.87E-11
ALActive cases/1,000 people5.40E-08NCActive cases/1,000 people7.49E-14
ARActive cases/1,000 people4.49E-14NDActive cases/1,000 people3.35E-13
AZActive cases/1,000 people7.12E-07NEUnemployment rate6.10E-10
CAActive cases/1,000 people2.29E-18NHImported COVID cases9.04E-32
COWork places1.48E-17NJUnemployment claims/1,000 people2.15E-07
CTTest capacity5.46E-25NMActive cases/1,000 people5.68E-13
DEResidential3.64E-15NVImported COVD cases2.36E-09
FLWork places9.60E-03NYResidential9.57E-05
GAActive cases/1,000 people2.76E-10OHActive cases/1,000 people6.42E-10
HIActive cases/1,000 people1.31E-08OKTests done/1,000 people4.14E-10
IAActive cases/1,000 people5.07E-14ORActive cases/1,000 people3.87E-12
IDImported COVD cases4.53E-10PAWork places4.21E-09
ILWork places5.82E-08RIImported COVD cases1.02E-27
INActive cases/1,000 people1.12E-09SCActive cases/1,000 people1.58E-04
KSTests done/1,000 people1.21E-32SDWork places3.89E-12
KYImported COVID cases4.30E-26TNActive cases/1,000 people3.45E-12
LA% Working from Home1.58E-19TXActive cases/1,000 people3.05E-06
MARetail3.20E-10UTActive cases/1,000 people2.04E-07
MDWork places6.35W-09VAActive cases/1,000 people2.08E-08
MEUnemployment rate7.15E-18VTUnemployment claims/1,000 people2.07E-32
MIWork places1.16E-10WAActive cases/1,000 people2.65E-21
MNActive cases/1,000 people9.01E-07WIActive cases/1,000 people5.19E-12
MOTests done/1,000 people2.42E-09WVActive cases/1,000 people1.53E-16
MSTests done/1,000 people2.16E-12WYActive cases/1,000 people2.13E-21

The most significant risk factor for the new cases of COVID-19 in each of 50 states in the US.

Table 5 summarized the most significant risk factor for the deaths from COVID-19 in each of the 50 states in the US. Active Cases/1,000 people, workplaces, residential, unemployment rate, imported COVID cases, unemployment claims/1,000 people, transit, test done/1,000 people, grocery, testing capacity, retail, percentage of change in consumption, percentage of working from home were the most significant risk factor for the deaths of COVID-19 in 17, 10, 4, 4, 3, 2, 2, 2, 1, 1, 1, 1 states, respectively. We also observed that the number of protest attendees showed mild significant evidence to cause increasing the number of new cases of COVID-19 in KY (p-value < 0.00012), KS (p-value < 0.00026), NH (p-value < 0.00108), MA (p-value < 0.0016) and TN (p-value < 0.0024) or to cause more deaths from COVID-19 in OR (p-value < 5.11 E-05), TX (p-value < 0.00017), ME (p-value < 0.00028), KS (p-value < 0.00061), MI (p-value < 0.0015), OH (p-value < 0.0021) and NC (p-value < 0.0023).

TABLE 5

StateRisk factorp-valueStateRisk factorp-value
AKActive cases/1,000 people3.95E-24MTActive cases/1,000 people4.41E-25
ALActive cases/1,000 people1.10E-18NCWork places8.95E-18
ARTests Done/1,000 people1.54E-30NDUnemployment rate1.04E-20
AZActive cases/1,000 people1.29E-27NEWork places2.62E-23
CAWork places1.54E-13NHImported COVID cases3.41E-28
COResidential6.44E-26NJResidential1.96E-09
CTActive cases/1,000 people3.12E-16NMUnemployment rate1.15E-16
DERetail1.35E-06NVActive cases/1,000 people1.45E-17
FLActive cases/1,000 people1.47E-15NY% Change in Consumption1.52E-05
GAWork places5.81E-17OHUnemployment rate2.67E-17
HIActive cases/1,000 people1.14E-20OKWork places1.89E-19
IAUnemployment rate8.06E-15OR% Working from Home4.50E-26
IDActive cases/1,000 people7.11E-18PAImported COVID cases1.11E-23
ILResidential1.34E-12RIImported COVID cases2.16E-22
INWork places3.39E-17SCActive cases/1,000 people6.75E-25
KSGrocery1.42E-40SDActive cases/1,000 people6.01E-24
KYWork places1.40E-24TNTest Done/1,000 people2.06E-24
LATransit2.38E-20TXActive cases/1,000 people2.59E-23
MAActive cases/1,000 people3.92E-09UTActive cases/1,000 people8.12E-29
MDGrocery7.60E-06VAWork places2.61E-16
MEResidential4.29E-25VTUnemployment claims/1,000 people1.23E-29
MIUnemployment claims/1,000 people2.07E-05WATransit3.83E-29
MNTesting capacity8.89E-11WIWork places4.77E-29
MDWork places7.85E-20WVActive cases/1,000 people7.81E-22
MSActive cases/1,000 people1.49E-19WYActive cases/1,000 people2.39E-28

The most significant risk factor for deaths of COVID-19 in each of 50 states in the US.

In some cases, two causal directions may occur. To examined whether the COVID-19 cases and deaths caused daily mobility and social distancing metrics to change, we summarized the values of statistics and p-values for testing COVID-19 new case and death potential causes of six Google mobility indexes across 50 states in the US in Supplementary Tables S5–7 and DS8, respectively. We did observe the opposite causal direction. Table 6 presented the smallest p-value across 50 states in the US for testing two causal directions: from Google mobility indexes to the number of new cases of COVID-19 and from the number of new cases of COVID-19 to the Google mobility indexes. We observed significant causation from the number of COVID-19to the Google mobility indexes in some states. However, the p-values for testing causation from the number of new cases of COVID-19 to the Google mobility indexes were much larger than that from the Google mobility indexes to the number of new cases of COVID-19. The causal pattern for the COVID-19 deaths was similar to the number of new cases of COVID-19.

TABLE 6

CauseEffectSmallest p-valueCauseEffectSmallest p-value
RetailCOVID-19 case7.67E-26COVID-19 caseRetail1.29E-04
Grocery3.70E-25Grocery4.85E-11
Parks3.87E-23Parks1.80E-12
Transit1.90E-27Transit1.29E-04
Workplaces1.07E-25Workplaces3.99E-06
Residential1.09E-25Residential2.99E-07

Lower bound of p-values for testing causation from Google mobility indexes to COVID-19 case and from COVID-19 case to Google mobility indexes.

To illustrate the causal relationships between the risk factors and the number of new cases and deaths from COVID-19, we plotted Figures 1 and 2. Figure 1 plotted the social distance index curves as a function of time from March 5, 2020 to August 25, 2020 in Florida (FL) and Rhode Island (RI). Figure 1 showed that the social distance index in FL was much higher than that in RI state, which resulted in the larger number of new cases of COVID-19 in FL than that in RI. Figure 2 showed the number of imported COVID-19 cases as a function of time from March 5, 2020 to August 25, 2020 in Maryland (MD) and Wyoming (WY). We observed a huge difference in the number of imported COVID-19 cases between MD and RI. The very low number of imported cases of COVID-19 in WY resulted in the very low number of deaths from COVID-19 in WY, while the high number of imported cases in MD state led to the increased deaths from COVID-19 in MD. These results were consistent with the finding in the literature. It was reported that strong interventions would substantially decrease the number of deaths (Davies et al. 2020; Gagnon et al. 2020).

FIGURE 1

FIGURE 2

The COVID-19 case curves as a function of time from March 5, 2020 to August 25, 2020 in FL, RI, MD and WY were plotted in Figure 3, respectively. The response of COVID-19 to public health intervention which were measured by Google mobility indexes were usually delayed.

FIGURE 3

Discussion

Causal inference for COVID-19 is essential for selecting and implementing public intervention measures and understanding the role of the demographics in curbing the spread and reducing the deaths from COVID-19. In this paper, we systematically addressed the issues in identifying causal risk factors and evaluating the causal effects of risk factors and intervention measures on the spread and deaths from COVID-19 in the US. Risk factors and intervention measures included scalar variables and temporal variables. The ANMs were used to test for causal relationships between scalar risk factors and the average number of new cases or deaths from COVID-19 in the US. Transmission of COVID-19 is a dynamic system. Many risk factors and intervention measures are temporal variables. The Granger Causality Test was used to reveal the causal relationships between the temporal risk factors and intervention measures, and the number of new cases or new deaths from COVID-19 across the 50 states in the US.

The demographic risk factors were the major part of the scalar risk factors in the causal analysis of COVID-19. We found that population density was the most significant causal factor of both new cases and death from COVID-19. Population density measured the average number of people per square kilometer living in a built-up area. Densely populated states generated conditions where COVID-19 can spread quickly and undetected in the densely populated areas and created high levels of vulnerability. The second significant demographic factor was percentage of males. Our data suggested that men were more vulnerable to COVID-19 than women. However, our analysis did not conclude that more men than women were dying from COVID-19.

We also discovered that more Black Americans were dying from COVID-19. The reasons for this were complex. Black Americans had higher rates of chronic disease conditions, including diabetes, heart disease, and lung disease, were poor and more easily exposed to the COVID-19, and lived in the cramped housing. Inequities in the social determinants of health affected mortality and morbidity of COVID-19 for Hispanic Americans with much milder significance.

We studied the causal effect of major public health interventions across the 50 states in the US. In the absence of centralized intervention measures and implementation of a timeline and presence of the complex dynamics of human mobility and the variable intensity of local outbreaks of COVID-19, evaluating the causal effect of public health intervention measures on COVID-19 transmission and deaths in the USA posed a great challenge. We used 6 Google mobility indexes and 12 daily metrics to measure the effects of COVID-19 spread and public health interventions on mobility and social distancing, derived from mobile device location data and COVID-19 case data, provided by the University of Maryland COVID-19 Impact Analysis Platform. These real time metrics capture the dynamics of social distancing. Granger causality tests were used to identify the causality relationships between time series metrics and time varying in the number of new cases or deaths from COVID-19. Although the risk factors differed by location, Active Cases/1,000 people were a significant risk factor for both number of new cases and deaths from COVID-19 in most states. The most popular intervention measure in the US was workplaces (mobility trends for places of work). Workplaces were the significant cause of the number new cases of COVID-19 in 44 states and significant cause of death in 49 states. Therefore, workplaces should be considered as a very important risk mitigation measure to reduce the number of new cases and deaths from COVID-19. Tests done/1,000 people was the second population intervention in the US. It was the significant cause of the new cases of COVID-19 in 46 states and significant cause of death in 47 states. Virus test results in quick case identification and isolation to contain COVID-19, and rapid treatment to reduce the number of deaths. Imported COVID cases were also a top significant risk factor for speeding the spread and increasing the deaths from COVID-19. Our results showed that the imported COVID case metric was the significant causal factor for the new cases in 46 states and the significant causal factor for the deaths in 47 states.

Our results showed that the high numbers of cases and deaths from COVID-19 were due to lacking strong interventions and high population density. We observed that no metrics showed significant evidence in mitigating the COVID-19 epidemic in FL and only a few metrics showed evidence in reducing the number of new cases of COVID-19 in AZ, NY and TX. Our results showed strong interventions were needed to contain COVID-19.

Although we tried to systematically and comprehensively analyze the data, this study has multiple limitations. First, we only analyzed the causal relationship between mobility patterns and the number of new cases or deaths and ignored the role of other potential mitigating factors (e.g., wearing face masks) that could also have contributed to the reduction of new cases or deaths from COVID-19. When data are available, more metrics should be included in the analysis.

Second, we have not addressed the confounding bias issue. When confounding is unknown, adjusting for confounding methods cannot be applied to eliminate confounding bias from the causal analysis. Unadjusted confounding bias will distort the inferred (true) causal relationship between the number of new cases or deaths from COVID-19, and metrics for social distancing when these two variables share common causes. This will have substantive implications for developing interventions to mitigate the spread of COVID-19 and reduce the deaths from COVID-19. However, removing confounding from causal analysis for COVID-19 is complicated and will be investigated in the future.

In summary, our analysis has provided information for both individuals and governments to plan future interventions on containing COVID-19 and reduction of deaths from COVID-19.

Data Availablity Statement

The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.

Funding

H-WD was partially supported by NIH Grants Nos. U19AG05537301 and R01AR069055. MX was partially supported by NIH Grants No. U19AG05537301.

Statements

Author contributions

ZL, data analysis; TX, data analysis; KZ, data acquisition and interpretation; H-WD, interpretation; EB, biological interpretation; MX, concept design, method development and write a manuscript.

Acknowledgments

The authors thank Sara Barton for editing the manuscript. The authors also thank the reviewers for their helpful comments.

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fams.2020.611805/full#supplementary-material.

References

Summary

Keywords

COVID-19, causal inference, time series, control of the spread, transmission dynamics, public health interventions

Citation

Li Z, Xu T, Zhang K, Deng H-W, Boerwinkle E and Xiong M (2021) Causal Analysis of Health Interventions and Environments for Influencing the Spread of COVID-19 in the United States of America. Front. Appl. Math. Stat. 6:611805. doi: 10.3389/fams.2020.611805

Received

29 September 2020

Accepted

03 December 2020

Published

25 January 2021

Volume

6 - 2020

Edited by

Zhanwei Du, University of Texas at Austin, United States

Reviewed by

Ting Tian, Sun Yat-sen University, China

Alessandro Rovetta, Mensana srls, Italy

Updates

Copyright

*Correspondence: Momiao Xiong,

This article was submitted to Mathematics of Computation and Data Science, a section of the journal Frontiers in Applied Mathematics and Statistics

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics