About this Event
2317 SPEEDWAY , Austin, Texas 78712
https://stat.utexas.edu/training/seminar-seriesRachel Nethery (Department of Biostatistics at the Harvard T.H. Chan School of Public Health)
Title: Bayesian and Machine Learning Approaches to Estimate Causal Effects in Environmental Health Applications
Abstract: I will discuss two projects on Bayesian and machine learning methods for causal inference applied to investigations of (1) the health impacts of exposure to natural gas infrastructure and (2) the causes of cancer clusters. In the first project, we seek to estimate the average causal effect of exposure to natural gas compressor stations on cancer mortality in the US. Because the data exhibit propensity score non-overlap (i.e. regions of poor support), estimation of population average causal effects requires reliance on model specifications. All existing methods to address non-overlap (e.g. trimming) change the estimand which can diminish the study's impact. We make two contributions on this topic. We first propose a data-driven definition of the overlap and non-overlap regions. Next, we develop a Bayesian machine learning method to estimate population average causal effects in the presence of non-overlap, which delegates the tasks of estimating causal effects in the overlap and non-overlap regions to two distinct models, suited to the degree of data support in each region.
In the second project, we propose a causal inference framework for cancer cluster investigations. These investigations arise when a community reports high cancer rates, often suspecting a relationship between the cancer and a hazardous exposure. Departments of health typically perform a standardized incidence ratio (SIR) analysis in response. This approach has several well-documented limitations. Assuming that a potentially hazardous exposure in the community is identified a priori, we introduce an estimand called the causal SIR (cSIR): the expected cancer incidence in the exposed population divided by the expected cancer incidence for the same population under the counterfactual scenario of no exposure. To estimate the cSIR we must (1) identify unexposed populations similar to the exposed one to inform estimation of the counterfactual and (2) resolve the spatial over-aggregation of publicly available cancer incidence data for these unexposed populations. We address the first challenge with matching and the second by developing a Bayesian model that borrows information from other sources to impute cancer incidence at the desired level of spatial aggregation.
0 people are interested in this event
User Activity
No recent activity