<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>censoring | Chen Xing</title>
    <link>https://chenxing.space/tag/censoring/</link>
      <atom:link href="https://chenxing.space/tag/censoring/index.xml" rel="self" type="application/rss+xml" />
    <description>censoring</description>
    <generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Tue, 17 Mar 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://chenxing.space/media/sharing.png</url>
      <title>censoring</title>
      <link>https://chenxing.space/tag/censoring/</link>
    </image>
    
    <item>
      <title>Notes on Doubly Robust Censoring Unbiased Transformation</title>
      <link>https://chenxing.space/blog/notes-on-doubly-robust-censoring-unbiased-transformation/</link>
      <pubDate>Tue, 17 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-doubly-robust-censoring-unbiased-transformation/</guid>
      <description>&lt;p&gt;Predicting outcomes with right-censored survival data forces a choice: do we model the outcome distribution, or the censoring mechanism? Classical transformations require us to commit to one. Rubin and van der Laan (2007) tell us we don&amp;rsquo;t have to. Their &lt;mark&gt;&lt;strong&gt;doubly robust censoring unbiased transformation&lt;/strong&gt;&lt;/mark&gt; fuses both approaches, remaining valid as long as &lt;em&gt;at least one&lt;/em&gt; of the two nuisance models is correctly specified. This post walks through the setup, the classical transformations, and how the doubly robust version combines them — drawing the analogy to AIPW along the way.&lt;/p&gt;
&lt;h2 id=&#34;1-setup&#34;&gt;1. Setup&lt;/h2&gt;
&lt;p&gt;We observe an i.i.d. sample  $\{O_i\}_{i=1}^n$  where each observation is&lt;/p&gt;
 $$O = \bigl(W,\; \Delta = \mathbf{1}(Y \le C),\; \tilde{Y} = Y \wedge C\bigr)$$ 
&lt;ul&gt;
&lt;li&gt;$W$: covariates&lt;/li&gt;
&lt;li&gt;$Y$: true (possibly unobserved) survival time&lt;/li&gt;
&lt;li&gt;$C$: random censoring time&lt;/li&gt;
&lt;li&gt;$\tilde{Y} = \min(Y, C)$ is what we actually record&lt;/li&gt;
&lt;li&gt;$\Delta = 1$ if the event is observed ($Y \le C$), and $\Delta = 0$ if the outcome is right-censored ($Y &amp;gt; C$)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Our goal&lt;/strong&gt; is to estimate the regression function&lt;/p&gt;
&lt;p&gt;$$m(w) = \mathbb{E}[Y \mid W = w]$$&lt;/p&gt;
&lt;p&gt;Two nuisance functions appear throughout:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$\bar{F}(\cdot \mid W)$: conditional survival function of the response $Y$ given $W$&lt;/li&gt;
&lt;li&gt;$\bar{G}(\cdot \mid W)$: conditional survival function of the censoring time $C$ given $W$&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We maintain the standard assumption $Y \indep C \mid W$.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;2-the-challenge-unidentifiability&#34;&gt;2. The Challenge: Unidentifiability&lt;/h2&gt;
&lt;p&gt;When the censoring time $C$ corresponds to a fixed study endpoint, the true response $Y$ may exceed it. Beyond that endpoint, &lt;strong&gt;nothing can be learned&lt;/strong&gt; about the tail of the survival distribution — making $m(W) = \mathbb{E}[Y \mid W]$ unidentifiable in general.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The fix:&lt;/strong&gt; truncate the response at a known horizon $\tau$,&lt;/p&gt;
&lt;p&gt;$$Y \longmapsto Y \wedge \tau = \min(Y, \tau)$$&lt;/p&gt;
&lt;p&gt;and estimate $w \mapsto \mathbb{E}[Y \wedge \tau \mid W = w]$ instead.&lt;/p&gt;
&lt;h3 id=&#34;the-surrogate-response-strategy&#34;&gt;The Surrogate-Response Strategy&lt;/h3&gt;
&lt;p&gt;The general approach to prediction with right-censored data is:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Replace&lt;/strong&gt; the possibly unavailable responses  $\{Y_i\}_{i=1}^n$  with surrogate values  $\{Y^*(O_i)\}_{i=1}^n$  using an imputation map $Y^*(\cdot)$ built from observed data.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Plug&lt;/strong&gt; the imputed dataset  $\{W_i, Y^*(O_i)\}_{i=1}^n$  into any standard regression algorithm.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The imputation map $Y^*(\cdot)$ is called a &lt;mark&gt;&lt;strong&gt;censoring unbiased transformation&lt;/strong&gt;&lt;/mark&gt; (Fan and Gijbels 1996) if it satisfies&lt;/p&gt;
&lt;p&gt;$$\mathbb{E}[Y^*(O) \mid W] = \mathbb{E}[Y \mid W] = m(W)$$&lt;/p&gt;
&lt;p&gt;That is, the surrogate is an unbiased proxy for the true (latent) response, conditional on covariates.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;3-two-classical-transformations&#34;&gt;3. Two Classical Transformations&lt;/h2&gt;
&lt;h3 id=&#34;a-the-buckleyjames-transformation-depends-on-barf&#34;&gt;A. The Buckley–James Transformation (depends on $\bar{F}$)&lt;/h3&gt;
&lt;p&gt;The Buckley–James transformation imputes a censored observation with its conditional mean given that it exceeds the censoring time:&lt;/p&gt;
&lt;p&gt;$$Y^*(O) = \Delta Y + (1-\Delta) Q_{\bar{F}}(W, C)$$&lt;/p&gt;
&lt;p&gt;where $\Delta = 1$ if $Y$ is observed and $\Delta = 0$ if right-censored, and&lt;/p&gt;
 $$Q_{\bar{F}}(w, y) = \mathbb{E}[Y \mid W = w,\; Y &gt; y] = \frac{1}{\bar{F}(y \mid W=w)} \int_y^{+\infty} u \; dF(u \mid W=w)$$ 
&lt;p&gt;Intuitively: if we observe the event, we keep $Y$; if censored at $C$, we impute with the expected remaining survival time above $C$.&lt;/p&gt;
&lt;p&gt;This transformation requires correctly estimating $\bar{F}(\cdot \mid W)$, i.e., the conditional distribution of the survival time.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id=&#34;b-the-ipcw-transformation-depends-on-barg&#34;&gt;B. The IPCW Transformation (depends on $\bar{G}$)&lt;/h3&gt;
&lt;p&gt;Inverse probability of censoring weighting (IPCW) up-weights the observed events to compensate for the censored ones:&lt;/p&gt;
&lt;p&gt;$$Y^*(O) = \frac{Y \Delta}{\bar{G}(Y \mid W)}$$&lt;/p&gt;
&lt;p&gt;This requires correctly estimating $\bar{G}(\cdot \mid W)$, the conditional survival function of the censoring time. It is the survival-analysis analogue of IPW in the causal inference literature.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;4-the-doubly-robust-censoring-unbiased-transformation&#34;&gt;4. The Doubly Robust Censoring Unbiased Transformation&lt;/h2&gt;
&lt;p&gt;The two classical transformations each stake everything on one nuisance model. The doubly robust approach combines them:&lt;/p&gt;

$$
Y^*(O) = \underbrace{\frac{Y\Delta}{\bar{G}(Y \mid W)}}_{\text{1st term}} + \underbrace{\frac{Q_{\bar{F}}(W,C)\,(1-\Delta)}{\bar{G}(C \mid W)}}_{\text{2nd term}} - \underbrace{\int_{-\infty}^{\tilde{Y}} \frac{Q_{\bar{F}}(W,c)}{\bar{G}^2(c \mid W)}\, dG(c \mid W)}_{\text{3rd term (correction)}}
$$

&lt;p&gt;where $\tilde{Y} = Y \wedge C = \min(Y,C)$ and $Q_{\bar{F}}(w, c) = \mathbb{E}[Y \mid W=w,; Y &amp;gt; c]$.&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$\mathbb{E}[Y^*(O) \mid W] = \mathbb{E}[Y \mid W]$ whenever either $\bar{F}(\cdot \mid W)$ or $\bar{G}(\cdot \mid W)$ is correctly specified.

  &lt;/div&gt;
&lt;/div&gt;
&lt;hr&gt;
&lt;h2 id=&#34;5-intuition-an-aipw-in-disguise&#34;&gt;5. Intuition: An AIPW in Disguise&lt;/h2&gt;
&lt;p&gt;The three-term structure has a clean interpretation. Recall that $Q_{\bar{F}}(W,C) = \mathbb{E}[Y \mid W, Y &amp;gt; C]$ is the &lt;strong&gt;outcome regression&lt;/strong&gt; for censored individuals.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Term&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;1st:&lt;/strong&gt; $Y\Delta \,/\, \bar{G}(Y\!\mid\! W)$&lt;/td&gt;
&lt;td&gt;For &lt;em&gt;observed&lt;/em&gt; outcomes — apply IPCW, the analogue of inverse propensity score weighting&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;2nd:&lt;/strong&gt; $Q_{\bar{F}}(W,C)(1-\Delta) \,/\, \bar{G}(C\!\mid\! W)$&lt;/td&gt;
&lt;td&gt;For &lt;em&gt;censored&lt;/em&gt; outcomes — impute with the outcome regression $Q_{\bar{F}}$, then apply an IPCW-style weight&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3rd:&lt;/strong&gt; $-\int Q_{\bar{F}} / \bar{G}^2 \; dG$&lt;/td&gt;
&lt;td&gt;Bias &lt;strong&gt;correction&lt;/strong&gt; term&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The first two terms together look exactly like &lt;strong&gt;IPW + imputation&lt;/strong&gt; — the two ingredients of AIPW in the standard (binary treatment) setting. The third term is the survival-analysis counterpart of the augmentation correction in AIPW: it removes the bias that accumulates when both nuisance models are slightly off.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;takeaway&#34;&gt;Takeaway&lt;/h2&gt;
&lt;p&gt;The doubly robust censoring unbiased transformation is a drop-in replacement for the classical Buckley–James or IPCW transformations. Once $Y^*(O_i)$ is computed for each observation, any off-the-shelf regression algorithm can be applied to the pairs $\{W_i, Y^*(O_i)\}$.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Rubin, Daniel and Mark J. van der Laan (2007), &amp;ldquo;A Doubly Robust Censoring Unbiased Transformation,&amp;rdquo; &lt;em&gt;The International Journal of Biostatistics&lt;/em&gt;, 3 (1).&lt;/p&gt;
&lt;p&gt;Fan, Jianqing and Irène Gijbels (1996), &lt;em&gt;Local Polynomial Modelling and Its Applications&lt;/em&gt;, Chapman &amp;amp; Hall.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on Propensity Score Methods</title>
      <link>https://chenxing.space/blog/notes-on-propensity-score/</link>
      <pubDate>Tue, 03 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-propensity-score/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Here are my notes on propensity scores, mainly from Prof. Ding&amp;rsquo;s textbook (2024).&lt;/p&gt;
&lt;p&gt;The traditional propensity score analysis workflow is shown in the image below, which I will not cover in detail. Instead, I will summarize the key theorems and results from Ding&amp;rsquo;s textbook.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/Image%20from%20Chap3.3_observational_PS,%20page%2016.png&#34; alt=&#34;Image from Chap3.3_observational_PS, page 16&#34; style=&#34;zoom:50%;&#34; /&gt;
  &lt;figcaption&gt;Figure 1: Traditional propensity score analysis workflow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I will also provide some connections with &lt;strong&gt;Riesz Representer (RR)&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Why connect with the Riesz Representer (RR)? The connection provides a powerful generalization of the foundational Rosenbaum-Rubin (1983) result.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Rosenbaum and Rubin showed that &lt;strong&gt;conditioning on the propensity score is sufficient for removing confounding bias&lt;/strong&gt; when estimating causal effects. The Riesz representer extends this principle: &lt;strong&gt;it suffices to regress on the Riesz representer&lt;/strong&gt; to obtain unbiased estimates of the average treatment effect.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;strong&gt;key insight&lt;/strong&gt; is that the Riesz representer, like the propensity score, serves as a sufficient statistic – it captures all the confounding information necessary for unbiased estimation of your target causal parameter.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;setting--notation&#34;&gt;Setting &amp;amp; Notation&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Binary treatment $Z$&lt;/li&gt;
&lt;li&gt;Potential outcomes $\{Y(0), Y(1)\}$  &lt;/li&gt;
&lt;li&gt;Propensity score: $\P(Z = 1 \mid X)$, where $X$ represents covariates&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Two approaches learning causal relationships:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Outcome process (via outcome regression)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Treatment assignment mechanism&lt;/strong&gt; (via propensity score)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The following summarizes the key theorems and results related to propensity scores from Prof. Ding&amp;rsquo;s textbook.&lt;/p&gt;
&lt;h2 id=&#34;1-the-propensity-score-as-a-markdimension-reductionmark-tool&#34;&gt;1. The propensity score as a &lt;mark&gt;dimension reduction&lt;/mark&gt; tool&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(pscore as dimension reduction tool)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$$
\text { If } Z \indep \{Y(1), Y(0)\} \mid X, \text { then } Z \indep \{Y(1), Y(0)\} \mid e(X) .
$$

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Covariates $X$ can be &lt;strong&gt;high dimensional&lt;/strong&gt;, but the propensity score, $e(X) \in \R$, is a  &lt;strong&gt;1-dimensional scalar&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;We can view the propensity score as a &lt;strong&gt;dimensional reduction&lt;/strong&gt; tool&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;2-propensity-score-stratification&#34;&gt;2. Propensity score stratification&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Idea&lt;/strong&gt;: Discretize the estimated propensity score by its $K$ quantiles:&lt;/p&gt;
 
$$
Z \indep \{Y(1), Y(0)\} \mid \hat{e}^{\prime}(X)=e_k \quad(k=1, \ldots, K) .
$$

&lt;p&gt;Estimate ATE within each subclass and then average by the block size&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Advantage&lt;/strong&gt;: The propensity score stratification estimator &lt;strong&gt;only requires the correct ordering&lt;/strong&gt; of the estimated propensity scores rather than their exact values, which makes it &lt;strong&gt;relatively robust&lt;/strong&gt; compared with other methods&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;3-propensity-score-weighting&#34;&gt;3. Propensity score weighting&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Invese propensity score weighting (IPW))&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
If $Z \indep \{Y(1), Y(0)\} \mid X$ and $ 0 &lt; e(X) &lt; 1$, then
$$E\{Y(1)\}=E\left\{\frac{Z Y}{e(X)}\right\}, \quad E\{Y(0)\}=E\left\{\frac{(1-Z) Y}{1-e(X)}\right\}$$
and 
$$
\begin{aligned}
\tau &amp;=E\{Y(1)-Y(0)\}\\
&amp;=E\left\{\frac{Z Y}{e(X)}-\frac{(1-Z) Y}{1-e(X)}\right\} \\
&amp;=E\left\{HY \right\}
\end{aligned}
$$
where
$H := \left[\frac{Z}{e(X)}-\frac{(1-Z) }{1-e(X)}\right]$ is called the &lt;strong&gt;Horvitz-Thompson transform&lt;/strong&gt;.

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Connection the Riesz Representer (RR)&lt;/p&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(RR in the case of ATE)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
    
    In the case of ATE, the Riesz Representer, $\alpha(Z, X)$, has the same form as above Horvitz-Thompson transform,
    $$
    \alpha(Z, X) = \left[\frac{Z}{e(X)}-\frac{(1-Z) }{1-e(X)}\right]
    $$
    
  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;31-estimation&#34;&gt;3.1 Estimation&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The sample version of IPW is called the &lt;strong&gt;Horvitz–Thompson (HT) estimator&lt;/strong&gt;,&lt;/p&gt;
 $$
\hat{\tau}^{\mathrm{ht}}=\frac{1}{n} \sum_{i=1}^n \frac{Z_i Y_i}{\hat{e}\left(X_i\right)}-\frac{1}{n} \sum_{i=1}^n \frac{\left(1-Z_i\right) Y_i}{1-\hat{e}\left(X_i\right)}
$$ 
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;HT estimator $\hat{\tau}^{\mathrm{ht}}$ has many problems&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Problem: lack of invariance&lt;/strong&gt;, i.e. if we replace $Y_i$ by $Y_i + c$, $\hat{\tau}^{\mathrm{ht}}$ changed because it depends on $c$. This is not reasonable.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Solution: normalizing the weights&lt;/strong&gt;&lt;/p&gt;
 $$
\hat{\tau}^{\text {hajek }}=\frac{\sum_{i=1}^n \frac{Z_i Y_i}{\hat{e}\left(X_i\right)}}{\sum_{i=1}^n \frac{Z_i}{\hat{e}\left(X_i\right)}}-\frac{\sum_{i=1}^n \frac{\left(1-Z_i\right) Y_i}{1-\hat{e}\left(X_i\right)}}{\sum_{i=1}^n \frac{1-Z_i}{1-\hat{e}\left(X_i\right)}} .
$$ 
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Hajek estimator is invariant to the location transformation&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;32-strong-overlap-condition&#34;&gt;3.2 Strong overlap condition&lt;/h3&gt;
&lt;p&gt;Many asymptotic analyses require a &lt;em&gt;strong overlap&lt;/em&gt; condition,&lt;/p&gt;
&lt;p&gt; $$
0&lt;\alpha_{\mathrm{L}} \leq e(X) \leq \alpha_{\mathrm{U}}&lt;1
$$ 
In practice,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Crump et al. (2009) suggested $α_L = 0.1$ and $α_U = 0.9$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Kurth et al. (2005) suggested $α_L = 0.05$ and $α_U = 0.95$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;4-balancing-property&#34;&gt;4. Balancing property&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-3&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 3&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(balancing property)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182217046.png&#34; alt=&#34;image-20250603182217046&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Conditional on $e(X)$, the treatment and the covariates are independent&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Within the same level of the propensity score, the covariate distributions are &lt;strong&gt;balanced&lt;/strong&gt; across the treatment and control groups&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Useful implication&lt;/strong&gt;: we can check whether the propensity score model is specified well enough to ensure the &lt;strong&gt;covariate balance&lt;/strong&gt; in the data&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;41-propensity-score-is-a-balancing-score&#34;&gt;4.1 Propensity score is a balancing score&lt;/h3&gt;







&lt;div class=&#34;math-environment definition&#34; id=&#34;definition-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Definition 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182507172.png&#34; alt=&#34;image-20250603182507172&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-4&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 4&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Propensity score is a balancing score)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182648222.png&#34; alt=&#34;image-20250603182648222&#34; style=&#34;zoom:50%;&#34; /&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182857253.png&#34; alt=&#34;image-20250603182857253&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;This is relevant in &lt;strong&gt;subgroup analysis&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The conditional independence in (11.5) &lt;mark&gt;ensures &lt;strong&gt;unconfoundedness&lt;/strong&gt; holds given the propensity score, within each level of $X_1$&lt;/mark&gt;. Therefore, we can perform the same analysis based on the propensity score, within each level of $X_1$, yielding estimates for two subgroup effects&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;5-doubly-robust-or-aipw&#34;&gt;5. Doubly Robust or AIPW&lt;/h2&gt;
&lt;p&gt;The following Theorem is summarized from Prof. Wager&amp;rsquo;s lecture notes (2024).&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-5&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 5&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(strong double robustness of AIPW estimator)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Define the outcome regression as $$
\mu_{(z)}(x)=\mathbb{E}\left[Y_i(z) \mid X_i=x\right],
$$
Define AIPW estimator as
$$
\begin{aligned}
\hat{\tau}_{A I P W} &amp; =\underbrace{\frac{1}{n} \sum_{i=1}\left(\hat{\mu}_{(1)}\left(X_i\right)-\hat{\mu}_{(0)}\left(X_i\right)\right)}_{\text { outcome regression estimator }} \\
&amp; +\underbrace{\frac{1}{n} \sum_{i=1}^n\left(\frac{Z_i}{\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(1)}\left(X_i\right)\right)-\frac{1-Z_i}{1-\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(0)}\left(X_i\right)\right)\right)}_{\text { applying IPW to the regression residuals }}
\end{aligned}
$$

If we use estimators $\hat{\mu}_{(z)}(x)$ and $\hat{e}(x)$ that are both consistent with root-mean squared error (RMSE) decaying faster than $n^{-\alpha_\mu}$ and $n^{-\alpha_e}$ respectively, and if furthermore $\alpha_\mu+\alpha_e \geq 1 / 2$, then

$$
\begin{aligned}
&amp; \sqrt{n}\left(\hat{\tau}_{A I P W}-\tau\right) \Rightarrow \mathcal{N}\left(0, V_{A I P W}\right) \\
&amp; V_{A I P W}=\operatorname{Var}\left[\tau\left(X_i\right)\right]+\mathbb{E}\left[\frac{\sigma_0^2\left(X_i\right)}{1-e\left(X_i\right)}\right]+\mathbb{E}\left[\frac{\sigma_1^2\left(X_i\right)}{e\left(X_i\right)}\right]
\end{aligned}
$$
where
$$
\sigma_{(z)}^2(x)=\operatorname{Var}\left[Y_i(z) \mid X_i=x\right]
$$


  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Check my previous post: &lt;a href=&#34;https://chenxing.space/blog/intuition-for-doubly-robust-estimator/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Intuition for Doubly Robust Estimator&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;AIPW provides a natural starting point for understanding Double Machine Learning&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key insight&lt;/strong&gt; of RR in DML framework: Leverage the Riesz Representer, a &amp;ldquo;generalized version of propensity score&amp;rdquo; to &amp;ldquo;correct the bias&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;6-other-estimands-related-to-ipw&#34;&gt;6. Other Estimands related to IPW&lt;/h2&gt;
&lt;p&gt;More general, Li et al. (2018a) gave a unified discussion of the causal estimands in observational studies.&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-6&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 6&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Ding (2024), Section 13.4)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602144337279.png&#34; alt=&#34;image-20250602144337279&#34; style=&#34;zoom:40%;&#34; /&gt;
&lt;br&gt;
Summary Table of common estimands: 
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602145602468.png&#34; alt=&#34;image-20250602145602468&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;This table provides us a good way to understand and remember IPW estimator for ATT&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How to remember $\tau^h$? Apply IPW on &lt;mark&gt;&amp;ldquo;pseudo outcome&amp;rdquo; $Yh(X)$ &lt;/mark&gt; then divide by $E(h(X))$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;When the parameter of interest is ATT, then $$E(h(X)) = E(e(X)) = E(E(Z \mid X)) = E(Z) = \P(Z = 1) = e$$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Use it to better understand IPW for ATT&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;7-propensity-score-in-regression&#34;&gt;7. Propensity Score in Regression&lt;/h2&gt;
&lt;h3 id=&#34;ps-as-a-covariate&#34;&gt;PS as a covariate&lt;/h3&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-7&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 7&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(regression with pscore as a covariate)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Under unconfoundedness, the coefficient of $Z$ in the population OLS fit of

$$
Y \sim 1+Z + e(X)
$$

equals &lt;mark&gt;$\tau_{\mathrm{O}}$&lt;/mark&gt;, $$
\tau_{\mathrm{O}}=\frac{E[e(X)\{1-e(X)\} \tau(X)]}{E[e(X)\{1-e(X)\}]}
,$$ which is the &lt;mark&gt;&lt;strong&gt;overlap-weighted average treatment effect&lt;/strong&gt;&lt;/mark&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Based on above Theorem, we also have:&lt;/p&gt;







&lt;div class=&#34;math-environment corollary&#34; id=&#34;corollary-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Corollary 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Under unconfoundedness, &lt;br&gt;

1. the coefficient of $Z$ in the population OLS fit of

$$
Y \sim 1+Z + e(X) + X
$$

also equals &lt;mark&gt;$\tau_{\mathrm{O}}$&lt;/mark&gt;, &lt;br&gt;&lt;/br&gt;

2. the coefficient of $Z-e(X)$ in the population OLS fit of

$$
Y \sim [Z - e(X)] \quad \text{or} \quad Y \sim 1 + [Z - e(X)]
$$

also equals &lt;mark&gt;$\tau_{\mathrm{O}}$&lt;/mark&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;ps-as-a-weight&#34;&gt;PS as a weight&lt;/h3&gt;
&lt;p&gt;There is a convenient way to obtain $\hat{\tau}^{\text{hajek}}$ based on WLS.&lt;/p&gt;







&lt;div class=&#34;math-environment proposition&#34; id=&#34;proposition-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Proposition 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(convenient to obtain $\hat{\tau}^{\text{hajek}}$ based on WLS)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250604101708438.png&#34; alt=&#34;image-20250604101708438&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Need to use bootstrap for standard error&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Why does the WLS give a consistent estimator for $\tau$ ?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In RCT with a constant propensity score, we can simply use the coefficient of $Z_i$ in the OLS fit of $Y_i$ on ( $1, Z_i$ ) to estimate $\tau$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In observational studies, we need to deal with the selection bias. The key idea is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;If we weight the treated units by $\frac{1}{e(X_i)}$ and the control units by $\frac{1}{1-e(X_i)}$, then both treated and control groups can represent the whole population&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Thus, &lt;strong&gt;by weighting, we effectively have a pseudo-randomized experiment&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(IPCW)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Inverse Probability of Censoring Weighting (IPCW) follows the same idea — it adjusts for censoring bias by reweighting observations based on their probability of being uncensored.

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Consequently, the difference between the weighted means is consistent for $\tau$. The numerical equivalence of $\hat{\tau}^{\text {hajek }}$ and WLS is not only a fun numerical fact itself but also useful for motivating more complex estimators with covariate adjustment&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Ding, Peng (2024), &lt;i&gt;A First Course in Causal Inference&lt;/i&gt;, CRC Press.&lt;/p&gt;
&lt;p&gt;Wager, S. (2024). Causal inference: A statistical learning approach. &lt;a href=&#34;https://web.stanford.edu/~swager/causal_inf_book.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://web.stanford.edu/~swager/causal_inf_book.pdf&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Adjust Censoring and Confounding Bias by IP Weighting</title>
      <link>https://chenxing.space/blog/how-to-adjust-censoring-bias-and-confounding-bias-with-ip-weights/</link>
      <pubDate>Tue, 01 Oct 2024 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/how-to-adjust-censoring-bias-and-confounding-bias-with-ip-weights/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;In the context of causal inference, adjusting for both censoring bias and confounding bias is crucial, particularly in survival analysis where right-censored data often complicates causal effect estimation. Right censoring occurs when the outcome of interest (e.g., time to an event) is not observed within the study period for some subjects, making it challenging to correctly assess the causal effect of a treatment. Moreover, confounding bias arises when treatment assignment is influenced by pre-treatment covariates, potentially leading to biased estimates of treatment effects if not properly accounted for.&lt;/p&gt;
&lt;p&gt;To address these issues, &lt;mark&gt;&lt;strong&gt;inverse probability weights (IPW)&lt;/strong&gt;&lt;/mark&gt; are commonly employed. IPW adjusts for confounding by reweighting observations based on their treatment probabilities given covariates. In addition, &lt;mark&gt;&lt;strong&gt;inverse probability of censoring weights (IPCW)&lt;/strong&gt;&lt;/mark&gt; further &lt;strong&gt;adjust for censoring by reweighting observations based on their probabilities of being uncensored&lt;/strong&gt;. Together, these techniques allow us to estimate causal effects in the presence of both confounding and censoring biases, providing more accurate insights into the relationships between treatments and outcomes.&lt;/p&gt;
&lt;p&gt;This post explains how to apply IP weights to adjust for these biases, with references to key resources.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Hernán and Robins’ &lt;em&gt;Causal Inference: What If&lt;/em&gt; (2020), especially &lt;em&gt;Chapter 8.5&lt;/em&gt; and &lt;em&gt;Chapter 12.6&lt;/em&gt;, covers censoring and missing data in detail&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Zubizarreta et al.’s &lt;em&gt;Handbook of Matching and Weighting Adjustments for Causal Inference&lt;/em&gt; (2023), Chapter 21, discusses treatment heterogeneity in survival outcomes. These resources form the basis for our explanation on using IP weights in survival analysis&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;ip-weighting&#34;&gt;IP Weighting&lt;/h2&gt;
&lt;p&gt;Imagine we want to estimate the causal effect of eco-friendly packaging on a product’s selling price. However, the selling price is right-censored — meaning we only observe the price for items that have been sold, while unsold items remain censored (i.e., their final selling price is unknown).&lt;/p&gt;
&lt;h3 id=&#34;setting&#34;&gt;Setting:&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Let $W \in \{0,1\}$ be a binary &lt;strong&gt;treatment&lt;/strong&gt; variable (e.g., $W= 1$, if the item has a eco-friendly package; $W = 0$, otherwise)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let $Y(w, s)$ be potential &lt;strong&gt;outcomes&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let $S = \textbf{1} \{C = 0\}$ be &lt;strong&gt;non-censoring indicator&lt;/strong&gt;, where $C \in \{0,1\}$ is the censoring indicator&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;e.g. $S = 1$, if the item is non-censored (in this case sold); $S = 0$, if the item is censored (in this case on-sale)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let $X, L \in \mathcal{X}$ be two sets of &lt;strong&gt;covariates&lt;/strong&gt; and $X \subset L$. Then there exist a function $f(\cdot)$ such that $X = f(L)$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Considering $Y(W = w, S = s)$, our analysis was necessarily restricted to uncensored individuals, i.e., those with $S = 1$, because those were the only ones with known values of the outcome $Y$. Thus, the causal effect of interest is:&lt;/p&gt;

$$
\tau = \mathbb{E}\{Y(w = 1, s = 1)\} - \mathbb{E}\{Y(w = 0, s= 1)\}
$$

&lt;h3 id=&#34;assumptions&#34;&gt;Assumptions:&lt;/h3&gt;
&lt;p&gt;Recall the three identification conditions in &lt;em&gt;Chapter 8.5&lt;/em&gt; (Hernán and Robins, 2020, p. 113-114). Note that, the book uses $Y, A, C, L$ to represent the outcome, treatment, censoring indicator and covariates respectively.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;First, the average outcome in the uncensored individuals must equal the unobserved average outcome in the censored individuals with the same values of $A$ and $L$. This provision will be satisfied &amp;hellip; if the variables in $A$ and $L$ are sufficient to block all backdoor paths between $C$ and $Y$.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Second, IP weighting requires that all conditional probabilities of being uncensored given A and the variables in L must be greater than zero.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;The third condition is consistency, including sufficiently well-deﬁned interventions. IP weighting is used to create a pseudo-population in which censoring $C$ has been abolished, and in which the effect of the treatment $A$ is the same as in the original population.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Following above three identification conditions, we make the following assumptions. Denote the &lt;strong&gt;propensity score&lt;/strong&gt; as $e(\cdot)$ and &lt;strong&gt;censoring score&lt;/strong&gt; as $c(\cdot)$,&lt;/p&gt;

$$
\begin{alignat}{2}
    &amp;\text{1. (Unconfoundedness):} \quad &amp;&amp; \{Y(w= 0, s = 1), Y(w= 1, s = 1)\} \perp \!\!\! \perp W \mid X \\[1.5em]
    &amp;\text{2. (Overlap):} \quad &amp;&amp; 0 &lt; e(X) := \mathbb{P}(W = 1 \mid X) &lt; 1 \\[1.5em]
    &amp;\text{3. (Ignorable censoring):} \quad &amp;&amp; \{Y(w= 0, s = 1), Y(w= 1, s = 1)\} \perp \!\!\! \perp S \mid (W, L) \\[1.5em]
    &amp;\text{4. (Positivity):} \quad &amp;&amp; c(W, L) := \mathbb{P}(S = 1 \mid (W, L)) &gt; 0
\end{alignat}
$$

&lt;h3 id=&#34;claim&#34;&gt;Claim:&lt;/h3&gt;

$$
\begin{align}
    \tau &amp;= \mathbb{E}\left[Y(w= 1, s = 1) - Y(w= 0, s = 1)\right] \\[1em]
    &amp;= \mathbb{E}\left[\frac{S W Y}{c(W, L) e(X)} - \frac{S (1-W) Y}{c(W, L) (1-e(X))}\right] \quad \tag{2}
\end{align}
$$

&lt;h3 id=&#34;proof&#34;&gt;Proof:&lt;/h3&gt;
&lt;p&gt;First, note that the eqn (2) is well defined because of A2 and A4.&lt;/p&gt;
&lt;p&gt;I only give the proof for,
$$\mathbb{E}\left[Y(w= 1, s = 1)\right] = \mathbb{E}\left[\frac{S W Y}{c(W, L) e(X)}\right],$$ as the $\mathbb{E}\left[Y(w= 0, s = 1)\right]$ part follows the same logic.&lt;/p&gt;
&lt;p&gt;Recall that,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;By our setting, $X = f(L)$ for some known function $f$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$c(W, L) := \mathbb{P}(S = 1 \mid W, L) =  \mathbb{E}\left[S \mid W,L \right]$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Then we have,&lt;/p&gt;

\begin{aligned}
    \mathbb{E}\left[\frac{S W Y}{c(W, L) e(X)}\right] 
    &amp;= \mathbb{E}\left\{\frac{W}{e(f(L))c(W,L))} \cdot \mathbb{E}\left[SY(w, s=1) \,\middle|\, W,L \right] \right\} &amp;&amp;\text{(by LIE)} \\[1.5em]
    &amp; = \mathbb{E}\left\{\frac{W}{e(f(L))c(W,L))} \cdot \mathbb{E}\left[S \,\middle|\, W,L \right] \cdot \mathbb{E}\left[Y(w, s=1) \,\middle|\, W,L \right] \right\} &amp;&amp;\text{(by A3)} \\[1.5em]
    &amp; = \mathbb{E}\left\{\frac{W}{e(f(L))} \cdot \mathbb{E}\left[Y(w, s=1) \,\middle|\, W,L \right] \right\}  \\[1.5em]
    &amp; = \mathbb{E}\left\{\mathbb{E}\left[\frac{W}{e(f(L))} \cdot Y(w=1, s=1) \,\middle|\, W,L \right] \right\} &amp;&amp;\text{(by SUTVA)}  \\[1.5em]
    &amp; = \mathbb{E}\left[\frac{W}{e(X)} \cdot Y(w=1, s=1) \right]  \\[1.5em]
    &amp; = \mathbb{E}\left[Y(w=1, s=1) \right] &amp;&amp;\text{(by A1 and IPW)}
\end{aligned}

&lt;p&gt;The last equation holds by the classical proof of the unbiasedness of inverse propensity score weighting (IPW) estimator. For more details, one can check Theorem 11.3 at page 158-159 in Peng Ding&amp;rsquo;s textbook (Ding, 2023).&lt;/p&gt;
&lt;p align=&#39;right&#39;&gt;Q.E.D.&lt;/p&gt;
&lt;h2 id=&#34;future-work&#34;&gt;Future Work&lt;/h2&gt;
&lt;p&gt;How to create a &lt;strong&gt;doubly robust&lt;/strong&gt; estimator based on above setting? One related literature is the causal survival forest model (Cui et al., 2023), which estimates heterogeneous treatment effects in time-to-event setting and obtains doubly robustness property. But sometimes we are more interested in the downstream outcomes (e.g. selling price) after the event (e.g. being sold). This motivates us to create a new estimator that is doubly robust&amp;hellip;&lt;/p&gt;
&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;p&gt;Hernán MA, Robins JM (2020). Causal Inference: What If. Boca Raton: Chapman &amp;amp; Hall/CRC.&lt;/p&gt;
&lt;p&gt;Zubizarreta, J. R., Stuart, E. A., Small, D. S., &amp;amp; Rosenbaum, P. R. (2023). &lt;i&gt;Handbook of Matching and Weighting Adjustments for Causal Inference&lt;/i&gt;. CRC Press.&lt;/p&gt;
&lt;p&gt;Cheng, C., Li, F., Thomas, L. E., &amp;amp; Li, F. (Frank). (2022). Addressing Extreme Propensity Scores in Estimating Counterfactual Survival Functions via the Overlap Weights. &lt;i&gt;American Journal of Epidemiology&lt;/i&gt;, &lt;i&gt;191&lt;/i&gt;(6), 1140–1151. &lt;a href=&#34;https://doi.org/10.1093/aje/kwac043&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1093/aje/kwac043&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=j1lFiviKcmM&amp;amp;t=1672s&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;YouTube tutorial about IPCW: &amp;ldquo;Survival Analysis, Censoring and Time Scales&amp;rdquo;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Cui, Y., Kosorok, M. R., Sverdrup, E., Wager, S., &amp;amp; Zhu, R. (2023). Estimating heterogeneous treatment effects with right-censored data via causal survival forests. &lt;i&gt;Journal of the Royal Statistical Society Series B: Statistical Methodology&lt;/i&gt;, &lt;i&gt;85&lt;/i&gt;(2), 179–211. &lt;a href=&#34;https://doi.org/10.1093/jrsssb/qkac001&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1093/jrsssb/qkac001&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Ding, Peng. “A First Course in Causal Inference.” arXiv, October 3, 2023. &lt;a href=&#34;https://doi.org/10.48550/arXiv.2305.18793&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.48550/arXiv.2305.18793&lt;/a&gt;.&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
