<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>propensity score | Chen Xing</title>
    <link>https://chenxing.space/tag/propensity-score/</link>
      <atom:link href="https://chenxing.space/tag/propensity-score/index.xml" rel="self" type="application/rss+xml" />
    <description>propensity score</description>
    <generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Wed, 22 Oct 2025 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://chenxing.space/media/sharing.png</url>
      <title>propensity score</title>
      <link>https://chenxing.space/tag/propensity-score/</link>
    </image>
    
    <item>
      <title>Balancing Weights for Causal Inference</title>
      <link>https://chenxing.space/blog/balancing-weights-for-causal-inference/</link>
      <pubDate>Wed, 22 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/balancing-weights-for-causal-inference/</guid>
      <description>&lt;h2 id=&#34;tldr&#34;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;Cohn et al. (2023) introduces the &lt;strong&gt;balancing approach&lt;/strong&gt; to weighting for causal inference in observational studies. &lt;mark&gt;Unlike traditional methods that model the propensity score directly, balancing weights are estimated by solving an optimization problem that directly targets covariate balance between treatment groups.&lt;/mark&gt; The authors demonstrate that this approach offers protection against model misspecification, connects naturally to bias-variance trade-offs, and can be augmented with outcome modeling for improved performance. Applied to the classic LaLonde job training data, balancing methods achieve better covariate balance than standard propensity score approaches while maintaining reasonable effective sample sizes.&lt;/p&gt;
&lt;h2 id=&#34;what-is-this-paper-about&#34;&gt;What is this paper about?&lt;/h2&gt;
&lt;p&gt;Covariate balance is fundamental to causal inference: randomized experiments achieve it by design, while observational studies must adjust for it. This chapter addresses a key challenge in observational causal inference—how to construct weights that remove confounding by balancing observed covariates between treated and control groups. The traditional modeling approach estimates propensity scores (the probability of treatment given covariates) and inverts them to create weights, but this relies heavily on correct model specification. When the propensity score model is wrong, the resulting weights may fail to balance covariates in the sample, leading to biased treatment effect estimates. &lt;mark&gt;The chapter explores an alternative: directly finding weights that achieve balance in the observed data, rather than first modeling the propensity score.&lt;/mark&gt;&lt;/p&gt;
&lt;h2 id=&#34;what-do-the-authors-do&#34;&gt;What do the authors do?&lt;/h2&gt;
&lt;p&gt;The authors formalize the balancing approach as an optimization problem that jointly minimizes covariate imbalance and weight dispersion (variance). They show how different choices of the &amp;ldquo;model class&amp;rdquo; M—the set of functions of covariates to balance—correspond to different assumptions about the outcome model and lead to different optimization formulations. Using the LaLonde dataset (a constructed observational study where the true treatment effect is known), they compare three designs: balancing main covariate terms only, balancing up to three-way interactions, and balancing an infinite-dimensional reproducing kernel Hilbert space (RKHS). For each design, they evaluate covariate balance using standardized mean differences, examine the effective sample size (a measure of weight dispersion), and explore the bias-variance trade-off by varying regularization parameters. The authors also demonstrate how balancing weights can be augmented with outcome regression to further reduce bias, and they establish asymptotic normality results for inference.&lt;/p&gt;
&lt;h2 id=&#34;why-is-this-important&#34;&gt;Why is this important?&lt;/h2&gt;
&lt;p&gt;This work matters because most observational studies include covariates in their analysis, yet practitioners often don&amp;rsquo;t carefully consider whether their weighting method actually achieves balance on the relevant covariate functions. The balancing approach makes covariate balance a first-order design criterion rather than a post-hoc diagnostic check. It reveals the implicit bias-variance trade-offs in weighting methods and shows that different assumptions about the outcome model (linear, interactive, nonparametric) lead to fundamentally different weighting strategies. &lt;mark&gt;The framework &lt;strong&gt;unifies&lt;/strong&gt; many existing methods (entropy balancing, stable weights, kernel balancing) under one optimization structure and clarifies the connection between the balancing and modeling approaches through Lagrangian duality.&lt;/mark&gt; Importantly, the chapter provides practical guidance on design choices—what to balance, how much dispersion to tolerate, whether to allow negative weights—that applied researchers face but often lack principled ways to resolve.&lt;/p&gt;
&lt;h2 id=&#34;who-should-care&#34;&gt;Who should care?&lt;/h2&gt;
&lt;p&gt;Applied researchers in economics, epidemiology, public policy, education, and medicine who use inverse propensity weighting or other covariate adjustment methods in observational studies. Methodologists working on causal inference, especially those developing new weighting estimators or studying properties of existing ones. Graduate students learning causal inference who need to understand the trade-offs between different adjustment strategies and the implicit assumptions behind common practices. Policy evaluators who must justify their modeling choices and demonstrate that their treatment effect estimates are robust to covariate imbalance. Anyone who has struggled with poor covariate balance after propensity score weighting or wondered how to choose between competing adjustment methods would benefit from this framework.&lt;/p&gt;
&lt;h2 id=&#34;do-we-have-code&#34;&gt;Do we have code?&lt;/h2&gt;
&lt;p&gt;The chapter does not provide standalone replication code or software packages. However, the authors note that many of the specific balancing methods discussed are available in existing R packages: entropy balancing in the &lt;code&gt;ebal&lt;/code&gt; package, covariate balancing propensity scores (CBPS) in the &lt;code&gt;CBPS&lt;/code&gt; package, and stable balancing weights in the &lt;code&gt;sbw&lt;/code&gt; package. The kernel balancing approach can be implemented using standard kernel methods in R or Python. The LaLonde dataset used throughout the examples is publicly available and widely used in the causal inference literature, making it straightforward to reproduce the analyses with these existing tools.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;In summary&lt;/strong&gt;, this chapter reframes propensity score weighting as a balance-optimization problem rather than a pure modeling exercise. By directly targeting the balancing property of inverse propensity weights, the approach offers robustness to propensity score misspecification while making explicit the bias-variance trade-offs inherent in any covariate adjustment. The LaLonde application demonstrates that balancing weights can substantially reduce covariate imbalance compared to standard methods, though at the cost of reduced effective sample size. The framework provides both theoretical insight (connecting balancing to dual regression, establishing asymptotic properties) and practical guidance (how to choose what to balance, when to augment with outcome modeling, whether to allow extrapolation) that fills an important gap in applied causal inference.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Cohn, Eric R., Eli Ben-Michael, Avi Feller, and José R. Zubizarreta (2023), “Balancing Weights for Causal Inference,” in Handbook of Matching and Weighting Adjustments for Causal Inference, Chapman and Hall/CRC. &lt;a href=&#34;https://www.taylorfrancis.com/chapters/edit/10.1201/9781003102670-16/balancing-weights-causal-inference-eric-cohn-eli-ben-michael-avi-feller-jos&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://www.taylorfrancis.com/chapters/edit/10.1201/9781003102670-16/balancing-weights-causal-inference-eric-cohn-eli-ben-michael-avi-feller-jos&lt;/a&gt;é-zubizarreta&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on Propensity Score Methods</title>
      <link>https://chenxing.space/blog/notes-on-propensity-score/</link>
      <pubDate>Tue, 03 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-propensity-score/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Here are my notes on propensity scores, mainly from Prof. Ding&amp;rsquo;s textbook (2024).&lt;/p&gt;
&lt;p&gt;The traditional propensity score analysis workflow is shown in the image below, which I will not cover in detail. Instead, I will summarize the key theorems and results from Ding&amp;rsquo;s textbook.&lt;/p&gt;
&lt;figure&gt;
  &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/Image%20from%20Chap3.3_observational_PS,%20page%2016.png&#34; alt=&#34;Image from Chap3.3_observational_PS, page 16&#34; style=&#34;zoom:50%;&#34; /&gt;
  &lt;figcaption&gt;Figure 1: Traditional propensity score analysis workflow&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I will also provide some connections with &lt;strong&gt;Riesz Representer (RR)&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Why connect with the Riesz Representer (RR)? The connection provides a powerful generalization of the foundational Rosenbaum-Rubin (1983) result.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Rosenbaum and Rubin showed that &lt;strong&gt;conditioning on the propensity score is sufficient for removing confounding bias&lt;/strong&gt; when estimating causal effects. The Riesz representer extends this principle: &lt;strong&gt;it suffices to regress on the Riesz representer&lt;/strong&gt; to obtain unbiased estimates of the average treatment effect.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;strong&gt;key insight&lt;/strong&gt; is that the Riesz representer, like the propensity score, serves as a sufficient statistic – it captures all the confounding information necessary for unbiased estimation of your target causal parameter.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;setting--notation&#34;&gt;Setting &amp;amp; Notation&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Binary treatment $Z$&lt;/li&gt;
&lt;li&gt;Potential outcomes $\{Y(0), Y(1)\}$  &lt;/li&gt;
&lt;li&gt;Propensity score: $\P(Z = 1 \mid X)$, where $X$ represents covariates&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Two approaches learning causal relationships:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Outcome process (via outcome regression)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Treatment assignment mechanism&lt;/strong&gt; (via propensity score)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The following summarizes the key theorems and results related to propensity scores from Prof. Ding&amp;rsquo;s textbook.&lt;/p&gt;
&lt;h2 id=&#34;1-the-propensity-score-as-a-markdimension-reductionmark-tool&#34;&gt;1. The propensity score as a &lt;mark&gt;dimension reduction&lt;/mark&gt; tool&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(pscore as dimension reduction tool)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$$
\text { If } Z \indep \{Y(1), Y(0)\} \mid X, \text { then } Z \indep \{Y(1), Y(0)\} \mid e(X) .
$$

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Covariates $X$ can be &lt;strong&gt;high dimensional&lt;/strong&gt;, but the propensity score, $e(X) \in \R$, is a  &lt;strong&gt;1-dimensional scalar&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;We can view the propensity score as a &lt;strong&gt;dimensional reduction&lt;/strong&gt; tool&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;2-propensity-score-stratification&#34;&gt;2. Propensity score stratification&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Idea&lt;/strong&gt;: Discretize the estimated propensity score by its $K$ quantiles:&lt;/p&gt;
 
$$
Z \indep \{Y(1), Y(0)\} \mid \hat{e}^{\prime}(X)=e_k \quad(k=1, \ldots, K) .
$$

&lt;p&gt;Estimate ATE within each subclass and then average by the block size&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Advantage&lt;/strong&gt;: The propensity score stratification estimator &lt;strong&gt;only requires the correct ordering&lt;/strong&gt; of the estimated propensity scores rather than their exact values, which makes it &lt;strong&gt;relatively robust&lt;/strong&gt; compared with other methods&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;3-propensity-score-weighting&#34;&gt;3. Propensity score weighting&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Invese propensity score weighting (IPW))&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
If $Z \indep \{Y(1), Y(0)\} \mid X$ and $ 0 &lt; e(X) &lt; 1$, then
$$E\{Y(1)\}=E\left\{\frac{Z Y}{e(X)}\right\}, \quad E\{Y(0)\}=E\left\{\frac{(1-Z) Y}{1-e(X)}\right\}$$
and 
$$
\begin{aligned}
\tau &amp;=E\{Y(1)-Y(0)\}\\
&amp;=E\left\{\frac{Z Y}{e(X)}-\frac{(1-Z) Y}{1-e(X)}\right\} \\
&amp;=E\left\{HY \right\}
\end{aligned}
$$
where
$H := \left[\frac{Z}{e(X)}-\frac{(1-Z) }{1-e(X)}\right]$ is called the &lt;strong&gt;Horvitz-Thompson transform&lt;/strong&gt;.

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Connection the Riesz Representer (RR)&lt;/p&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(RR in the case of ATE)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
    
    In the case of ATE, the Riesz Representer, $\alpha(Z, X)$, has the same form as above Horvitz-Thompson transform,
    $$
    \alpha(Z, X) = \left[\frac{Z}{e(X)}-\frac{(1-Z) }{1-e(X)}\right]
    $$
    
  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;31-estimation&#34;&gt;3.1 Estimation&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The sample version of IPW is called the &lt;strong&gt;Horvitz–Thompson (HT) estimator&lt;/strong&gt;,&lt;/p&gt;
 $$
\hat{\tau}^{\mathrm{ht}}=\frac{1}{n} \sum_{i=1}^n \frac{Z_i Y_i}{\hat{e}\left(X_i\right)}-\frac{1}{n} \sum_{i=1}^n \frac{\left(1-Z_i\right) Y_i}{1-\hat{e}\left(X_i\right)}
$$ 
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;HT estimator $\hat{\tau}^{\mathrm{ht}}$ has many problems&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Problem: lack of invariance&lt;/strong&gt;, i.e. if we replace $Y_i$ by $Y_i + c$, $\hat{\tau}^{\mathrm{ht}}$ changed because it depends on $c$. This is not reasonable.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Solution: normalizing the weights&lt;/strong&gt;&lt;/p&gt;
 $$
\hat{\tau}^{\text {hajek }}=\frac{\sum_{i=1}^n \frac{Z_i Y_i}{\hat{e}\left(X_i\right)}}{\sum_{i=1}^n \frac{Z_i}{\hat{e}\left(X_i\right)}}-\frac{\sum_{i=1}^n \frac{\left(1-Z_i\right) Y_i}{1-\hat{e}\left(X_i\right)}}{\sum_{i=1}^n \frac{1-Z_i}{1-\hat{e}\left(X_i\right)}} .
$$ 
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Hajek estimator is invariant to the location transformation&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;32-strong-overlap-condition&#34;&gt;3.2 Strong overlap condition&lt;/h3&gt;
&lt;p&gt;Many asymptotic analyses require a &lt;em&gt;strong overlap&lt;/em&gt; condition,&lt;/p&gt;
&lt;p&gt; $$
0&lt;\alpha_{\mathrm{L}} \leq e(X) \leq \alpha_{\mathrm{U}}&lt;1
$$ 
In practice,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Crump et al. (2009) suggested $α_L = 0.1$ and $α_U = 0.9$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Kurth et al. (2005) suggested $α_L = 0.05$ and $α_U = 0.95$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;4-balancing-property&#34;&gt;4. Balancing property&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-3&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 3&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(balancing property)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182217046.png&#34; alt=&#34;image-20250603182217046&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Conditional on $e(X)$, the treatment and the covariates are independent&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Within the same level of the propensity score, the covariate distributions are &lt;strong&gt;balanced&lt;/strong&gt; across the treatment and control groups&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Useful implication&lt;/strong&gt;: we can check whether the propensity score model is specified well enough to ensure the &lt;strong&gt;covariate balance&lt;/strong&gt; in the data&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;41-propensity-score-is-a-balancing-score&#34;&gt;4.1 Propensity score is a balancing score&lt;/h3&gt;







&lt;div class=&#34;math-environment definition&#34; id=&#34;definition-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Definition 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182507172.png&#34; alt=&#34;image-20250603182507172&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-4&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 4&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Propensity score is a balancing score)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182648222.png&#34; alt=&#34;image-20250603182648222&#34; style=&#34;zoom:50%;&#34; /&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250603182857253.png&#34; alt=&#34;image-20250603182857253&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;This is relevant in &lt;strong&gt;subgroup analysis&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The conditional independence in (11.5) &lt;mark&gt;ensures &lt;strong&gt;unconfoundedness&lt;/strong&gt; holds given the propensity score, within each level of $X_1$&lt;/mark&gt;. Therefore, we can perform the same analysis based on the propensity score, within each level of $X_1$, yielding estimates for two subgroup effects&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;5-doubly-robust-or-aipw&#34;&gt;5. Doubly Robust or AIPW&lt;/h2&gt;
&lt;p&gt;The following Theorem is summarized from Prof. Wager&amp;rsquo;s lecture notes (2024).&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-5&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 5&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(strong double robustness of AIPW estimator)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Define the outcome regression as $$
\mu_{(z)}(x)=\mathbb{E}\left[Y_i(z) \mid X_i=x\right],
$$
Define AIPW estimator as
$$
\begin{aligned}
\hat{\tau}_{A I P W} &amp; =\underbrace{\frac{1}{n} \sum_{i=1}\left(\hat{\mu}_{(1)}\left(X_i\right)-\hat{\mu}_{(0)}\left(X_i\right)\right)}_{\text { outcome regression estimator }} \\
&amp; +\underbrace{\frac{1}{n} \sum_{i=1}^n\left(\frac{Z_i}{\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(1)}\left(X_i\right)\right)-\frac{1-Z_i}{1-\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(0)}\left(X_i\right)\right)\right)}_{\text { applying IPW to the regression residuals }}
\end{aligned}
$$

If we use estimators $\hat{\mu}_{(z)}(x)$ and $\hat{e}(x)$ that are both consistent with root-mean squared error (RMSE) decaying faster than $n^{-\alpha_\mu}$ and $n^{-\alpha_e}$ respectively, and if furthermore $\alpha_\mu+\alpha_e \geq 1 / 2$, then

$$
\begin{aligned}
&amp; \sqrt{n}\left(\hat{\tau}_{A I P W}-\tau\right) \Rightarrow \mathcal{N}\left(0, V_{A I P W}\right) \\
&amp; V_{A I P W}=\operatorname{Var}\left[\tau\left(X_i\right)\right]+\mathbb{E}\left[\frac{\sigma_0^2\left(X_i\right)}{1-e\left(X_i\right)}\right]+\mathbb{E}\left[\frac{\sigma_1^2\left(X_i\right)}{e\left(X_i\right)}\right]
\end{aligned}
$$
where
$$
\sigma_{(z)}^2(x)=\operatorname{Var}\left[Y_i(z) \mid X_i=x\right]
$$


  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Check my previous post: &lt;a href=&#34;https://chenxing.space/blog/intuition-for-doubly-robust-estimator/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Intuition for Doubly Robust Estimator&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;AIPW provides a natural starting point for understanding Double Machine Learning&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key insight&lt;/strong&gt; of RR in DML framework: Leverage the Riesz Representer, a &amp;ldquo;generalized version of propensity score&amp;rdquo; to &amp;ldquo;correct the bias&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;6-other-estimands-related-to-ipw&#34;&gt;6. Other Estimands related to IPW&lt;/h2&gt;
&lt;p&gt;More general, Li et al. (2018a) gave a unified discussion of the causal estimands in observational studies.&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-6&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 6&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Ding (2024), Section 13.4)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602144337279.png&#34; alt=&#34;image-20250602144337279&#34; style=&#34;zoom:40%;&#34; /&gt;
&lt;br&gt;
Summary Table of common estimands: 
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602145602468.png&#34; alt=&#34;image-20250602145602468&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;This table provides us a good way to understand and remember IPW estimator for ATT&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How to remember $\tau^h$? Apply IPW on &lt;mark&gt;&amp;ldquo;pseudo outcome&amp;rdquo; $Yh(X)$ &lt;/mark&gt; then divide by $E(h(X))$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;When the parameter of interest is ATT, then $$E(h(X)) = E(e(X)) = E(E(Z \mid X)) = E(Z) = \P(Z = 1) = e$$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Use it to better understand IPW for ATT&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;7-propensity-score-in-regression&#34;&gt;7. Propensity Score in Regression&lt;/h2&gt;
&lt;h3 id=&#34;ps-as-a-covariate&#34;&gt;PS as a covariate&lt;/h3&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-7&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 7&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(regression with pscore as a covariate)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Under unconfoundedness, the coefficient of $Z$ in the population OLS fit of

$$
Y \sim 1+Z + e(X)
$$

equals &lt;mark&gt;$\tau_{\mathrm{O}}$&lt;/mark&gt;, $$
\tau_{\mathrm{O}}=\frac{E[e(X)\{1-e(X)\} \tau(X)]}{E[e(X)\{1-e(X)\}]}
,$$ which is the &lt;mark&gt;&lt;strong&gt;overlap-weighted average treatment effect&lt;/strong&gt;&lt;/mark&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Based on above Theorem, we also have:&lt;/p&gt;







&lt;div class=&#34;math-environment corollary&#34; id=&#34;corollary-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Corollary 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Under unconfoundedness, &lt;br&gt;

1. the coefficient of $Z$ in the population OLS fit of

$$
Y \sim 1+Z + e(X) + X
$$

also equals &lt;mark&gt;$\tau_{\mathrm{O}}$&lt;/mark&gt;, &lt;br&gt;&lt;/br&gt;

2. the coefficient of $Z-e(X)$ in the population OLS fit of

$$
Y \sim [Z - e(X)] \quad \text{or} \quad Y \sim 1 + [Z - e(X)]
$$

also equals &lt;mark&gt;$\tau_{\mathrm{O}}$&lt;/mark&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;ps-as-a-weight&#34;&gt;PS as a weight&lt;/h3&gt;
&lt;p&gt;There is a convenient way to obtain $\hat{\tau}^{\text{hajek}}$ based on WLS.&lt;/p&gt;







&lt;div class=&#34;math-environment proposition&#34; id=&#34;proposition-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Proposition 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(convenient to obtain $\hat{\tau}^{\text{hajek}}$ based on WLS)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250604101708438.png&#34; alt=&#34;image-20250604101708438&#34; style=&#34;zoom:50%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Need to use bootstrap for standard error&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Why does the WLS give a consistent estimator for $\tau$ ?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In RCT with a constant propensity score, we can simply use the coefficient of $Z_i$ in the OLS fit of $Y_i$ on ( $1, Z_i$ ) to estimate $\tau$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In observational studies, we need to deal with the selection bias. The key idea is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;If we weight the treated units by $\frac{1}{e(X_i)}$ and the control units by $\frac{1}{1-e(X_i)}$, then both treated and control groups can represent the whole population&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Thus, &lt;strong&gt;by weighting, we effectively have a pseudo-randomized experiment&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(IPCW)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Inverse Probability of Censoring Weighting (IPCW) follows the same idea — it adjusts for censoring bias by reweighting observations based on their probability of being uncensored.

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Consequently, the difference between the weighted means is consistent for $\tau$. The numerical equivalence of $\hat{\tau}^{\text {hajek }}$ and WLS is not only a fun numerical fact itself but also useful for motivating more complex estimators with covariate adjustment&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Ding, Peng (2024), &lt;i&gt;A First Course in Causal Inference&lt;/i&gt;, CRC Press.&lt;/p&gt;
&lt;p&gt;Wager, S. (2024). Causal inference: A statistical learning approach. &lt;a href=&#34;https://web.stanford.edu/~swager/causal_inf_book.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://web.stanford.edu/~swager/causal_inf_book.pdf&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on causal inference with No Overlap – Regression Discontinuity</title>
      <link>https://chenxing.space/blog/notes-on-causal-inference-with-no-overlap-regression-discontinuity/</link>
      <pubDate>Sat, 31 May 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-causal-inference-with-no-overlap-regression-discontinuity/</guid>
      <description>&lt;p&gt;Here is my notes on regression discontinuity from Prof. Ding&amp;rsquo;s textbook (2024) and Prof. Imai&amp;rsquo;s lecture notes.&lt;/p&gt;
&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;p&gt;We often cannot run a randomized experiment and have to use/design observational studies to find a setting where credible causal inference is possible.&lt;/p&gt;
&lt;p&gt;The key is the knowledge of &lt;strong&gt;treatment assignment mechanism&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Regression discontinuity design&lt;/strong&gt; (RD Design):&lt;/p&gt;
&lt;p&gt;RD Design is a simple and widely used &lt;strong&gt;quasi-experimental&lt;/strong&gt; design. The term “quasi experimental” is to emphasize that these approaches are still framed using concepts from randomized experiments but require econometric innovations to compensate for the lack of random treatment assignment.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Sharp RD Design&lt;/em&gt;: treatment assignment is based on a &lt;strong&gt;deterministic&lt;/strong&gt; rule (i.e. we have full knowledge of how treatment is assigned)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Fuzzy RD Design&lt;/em&gt;: &lt;strong&gt;encouragement to receive&lt;/strong&gt; treatment is based on a deterministic rule&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;setting&#34;&gt;Setting&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Binary treatment $Z\in \{0,1\}$  &lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Potential outcomes $\{Y(0), Y(1)\}$  &lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;There is a &lt;strong&gt;running variable&lt;/strong&gt; $X \in \R$ such that $Z=I\left(X \geq x_0\right)$, where $x_0$ is a pre-determined threshold. Note that, the treatment assignment is deterministic&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The unconfoundedness assumption holds automatically  $$
Z \indep \{Y(1), Y(0)\} \mid X
$$ &lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The overlap assumption does not hold&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;

$$
e(X)=\operatorname{pr}(Z=1 \mid X)=1\left(X \geq x_0\right) = \text{1 or 0}
$$

&lt;h2 id=&#34;identification&#34;&gt;Identification&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;RD can identify a &lt;strong&gt;local average causal effect&lt;/strong&gt; at the cutoff point $x_0$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Estimand:&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
  
$$
\tau\left(x_0\right)=E\left\{Y(1)-Y(0) \mid X=x_0\right\} .
$$

&lt;ul&gt;
&lt;li&gt;






&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(continuity assumption)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
1. $E\{Y(1) \mid X=x\}$ is continuous from the right at $x_0$ &lt;br&gt;
2. $E\{Y(0) \mid X=x\}$ is continuous from the left at $x_0$

  &lt;/div&gt;
&lt;/div&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250531113432572.png&#34; alt=&#34;image-20250531113432572&#34; style=&#34;zoom:20%;&#34; /&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;We have  $$
\begin{align}
E\left\{Y(1) \mid X=x_0\right\} &amp; =\lim _{\varepsilon \rightarrow 0+} E\left\{Y(1) \mid X=x_0+\varepsilon\right\} \tag{continuity} \\
&amp; =\lim _{\varepsilon \rightarrow 0+} E\left\{Y(1) \mid Z=1, X=x_0+\varepsilon\right\} \tag{def of Z}\\
&amp; =\lim _{\varepsilon \rightarrow 0+} E\left(Y \mid Z=1, X=x_0+\varepsilon\right),
\end{align}
$$  Similarly, $$
E\left\{Y(0) \mid X=x_0\right\}=\lim _{\varepsilon \rightarrow 0+} E\left(Y \mid Z=0, X=x_0-\varepsilon\right)
$$  So the local average causal effect at $x_0$ can be identified by the difference of the two limits&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Advantage: internal validity&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Disadvantage: external validity&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;key-theorem&#34;&gt;Key Theorem&lt;/h2&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Assume that the treatment is determined by $Z=I\left(X \geq x_0\right)$ where $x_0$ is a predetermined threshold. Assume that $E\{Y(1) \mid X=x\}$ is continuous from the right at $x_0$ and $E\{Y(0) \mid X=x\}$ is continuous from the left at $x_0$. Then the local average treatment effect at $X=x_0$ is identified by

$$
\tau\left(x_0\right)=\lim _{\varepsilon \rightarrow 0+} E\left(Y \mid Z=1, X=x_0+\varepsilon\right)-\lim _{\varepsilon \rightarrow 0+} E\left(Y \mid Z=0, X=x_0-\varepsilon\right)
$$

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;$\tau\left(x_0\right)$ is nonparametrically identified.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Ding, P. (2024). A First Course in Causal Inference. CRC Press.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://imai.fas.harvard.edu/teaching/files/regression_discontinuity.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Lecture notes: Regression Discontinuity Design&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on Callaway &amp; Sant’Anna (2021) – Staggered Adoption DiD</title>
      <link>https://chenxing.space/blog/notes-on-callaway-sant-anna-2021-staggered-adoption-did/</link>
      <pubDate>Tue, 27 May 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-callaway-sant-anna-2021-staggered-adoption-did/</guid>
      <description>&lt;h2 id=&#34;0-motivation&#34;&gt;0. Motivation&lt;/h2&gt;
&lt;p&gt;Staggered‐adoption policies break the canonical &lt;strong&gt;two-period / two-group&lt;/strong&gt; DiD model. It has been shown that the traditional two-way fixed-effects (TWFE) regression can assign &lt;strong&gt;negative weights&lt;/strong&gt; to treatment effects, thereby obscuring their dynamic and heterogeneous patterns.&lt;/p&gt;
&lt;p&gt;Callaway &amp;amp; Sant’Anna (2021) propose a &lt;strong&gt;divide-and-conquer&lt;/strong&gt; strategy:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Divide&lt;/strong&gt; the messy staggered panel into many honest $2\times2$ DiDs&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Conquer&lt;/strong&gt; by estimating each &amp;ldquo;little&amp;rdquo; DiD under familiar assumptions, then combine them with user-chosen weights to answer specific questions&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The key takeaway:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Use $\operatorname{ATT}(g, t)$ as a building block so we can transparently see how things are constructed&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Many different aggregation schemes are possible: they deliver different parameters&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Can allow for covariates via regressions adjustments, IPW, and DR.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;1-setup&#34;&gt;1. Setup&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Data structure&lt;/strong&gt;: Panel of units $i$ over time $t = 1,\dots,T$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let $D_{it}$ be a binary variable. $D_{it} = 1$ if unit $i$ is treated in period $t$; $D_{it} = 0$ otherwise&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Cohorts&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Define $G$ as the time period when a unit &lt;strong&gt;first becomes treated&lt;/strong&gt;. For all units that eventually get treated, $G$ defines which &amp;ldquo;group&amp;rdquo; they belong to&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Define $G_{g}$ as a binary variable. &lt;mark&gt;$G_{g} = 1$ if a unit is first treated in period $g$&lt;/mark&gt; (i.e.  $G_{i,g} = \1\{G_i = g\}$ )&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Define $G=\infty$ as &amp;ldquo;never treated&amp;rdquo; group.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Potential outcomes&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$Y_{it}(g)$: outcome at time $t$ if first treated in period $g$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$Y_{it}(\infty)$: outcome if never treated.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Cohort-time ATT&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Assume $\mathrm{iid}$, drop unit index $i$. A parameter of interest that has clear interpretation is the $\operatorname{ATT}(g, t)$: $$
\operatorname{ATT}(g, t)=\mathbb{E}\left[Y_t(g)-Y_t(\infty) \mid G_g=1\right], \text { for } t \geq g .
$$ This defines one &amp;ldquo;clean&amp;rdquo; $2\times2$ DiD for each pair $(g,t)$.&lt;/p&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Why focus on time after treatment starting period, $t \ge g$? Because we need the &lt;mark&gt;no anticipation&lt;/mark&gt; assumption. Before treatment taking place, $t &lt; g$, there is no treatment effect.

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;2-key-assumptions&#34;&gt;2. Key Assumptions&lt;/h2&gt;
&lt;p&gt;Given that we never observe $Y(\infty)$ in post-treatment periods among units that have been treated, we need to make assumptions to identify $\operatorname{ATT}(g, t)$.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(No anticipation)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$$
Y_{i t}(g)=Y_{i t}\left(g^{\prime}\right) \quad \forall \ i \text{, } t&lt;\min \left\{g, g^{\prime}\right\}
$$
Treatment cannot affect pre-treatment outcomes.

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;The no anticipation assumption has exactly the same content as in the $2\times2$ case.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Parallel trends based on &amp;#39;never-treated&amp;#39; control)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    


For each $t \in\{2, \ldots, T\}, g \in \mathcal{G}$ such that $\textcolor{red}{t \geq g}$,

$$
\mathbb{E}\left[Y_t(\infty)-Y_{t-1}(\infty) \mid G_g=1\right]=\mathbb{E}\left[Y_t(\infty)-Y_{t-1}(\infty) \mid G=\infty\right], \tag{1}
$$

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Note that, (1) is equivalent to the following: for  $t \in\{2, \ldots, T\}, g \in \mathcal{G}, t \geq g$ ,&lt;/p&gt;
&lt;p&gt;$$
\mathbb{E}\left[Y_t(\infty)-Y_{g-1}(\infty) \mid G_g=1\right]=\mathbb{E}\left[Y_t(\infty)-Y_{g-1}(\infty) \mid G=\infty\right], \tag{1&amp;rsquo;}
$$&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-3&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 3&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Parallel trends based on &amp;#39;Not-Yet-Treated&amp;#39; groups)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
For each $(s, t) \in\{2, \ldots, T\} \times\{2, \ldots, T\}, g \in \mathcal{G}$ such that $\textcolor{red}{t \geq g, s \geq t}$

$$
\mathbb{E}\left[Y_t(\infty)-Y_{t-1}(\infty) \mid G_g=1\right]=\mathbb{E}\left[Y_t(\infty)-Y_{t-1}(\infty) \mid D_s=0, G_g=0\right], \tag{2}
$$


  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Similarly, (2) is equivalent to (2&amp;rsquo;), that is, changing $Y_{t-1}(\infty)$ to $Y_{g-1(\infty)}$ in (2).&lt;/p&gt;
&lt;h2 id=&#34;3-identification-long-difference-estimands&#34;&gt;3. Identification: Long-Difference Estimands&lt;/h2&gt;
&lt;p&gt;Under no anticipation and one of the parallel-trends assumptions, each $\mathrm{ATT}(g,t)$ equals a simple &lt;mark&gt;&lt;strong&gt;long-difference DID&lt;/strong&gt;&lt;/mark&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Using never-treated&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
 $$
\operatorname{ATT}^{\text {never}}(g, t)=\mathbb{E}\left[Y_{t}-Y_{g-1} \mid G_g=1 \right]-\mathbb{E}\left[Y_{ t}-Y_{g-1} \mid G=\infty\right]
$$ 
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Using not-yet-treated&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
 $$
\operatorname{ATT}^{\text {not-yet}}(g, t)=\mathbb{E}\left[Y_{t}-Y_{g-1} \mid G_g=1 \right]-\mathbb{E}\left[Y_{ t}-Y_{g-1} \mid D_t = 0, G=\infty\right]
$$ 
&lt;figure style=&#34;text-align: center;&#34;&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250527100613626.png&#34; alt=&#34;image-20250527100613626&#34; style=&#34;zoom:50%;&#34; caption=&#34;sdf&#34;/&gt;
&lt;figcaption&gt;Why it’s called &lt;strong&gt;&#34;long difference&#34;&lt;/strong&gt;? Longer Time Span! &lt;/figcaption&gt;
&lt;/figure&gt;







&lt;div class=&#34;math-environment remark&#34; id=&#34;remark-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Remark 2&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Q: Why it’s called &lt;strong&gt;&#34;long difference&#34;&lt;/strong&gt;? &lt;br&gt;&lt;/br&gt;

A: Longer Time Span! Check my plot above. The &#34;long difference&#34; refers to the fact that the comparison often spans from a pre-treatment period (i.e. $g-1$) to a later period (post-treatment, $t \ge g$), potentially covering multiple time periods. This contrasts with &#34;short differences,&#34; which might involve comparing outcomes in consecutive periods or shorter time windows.&lt;br&gt;&lt;/br&gt;

For each treated cohort, the method computes the difference in outcomes between the pre-treatment period and a specific post-treatment period, potentially far apart in time. This extended gap emphasizes the &#34;long&#34; aspect, as it captures the cumulative effect of the treatment over time.



  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Moreover, with covariates, one can form &lt;strong&gt;doubly-robust&lt;/strong&gt; estimators that combine generalized propensity scores $p_g(X_i)$ and outcome models $m_{g,t}(X_i)$. At high level, the form of this estimator is identical to AIPW–ATT in cross–section, but we need to replace &amp;ldquo;levels&amp;rdquo; (e.g. $Y$) with &amp;ldquo;changes&amp;rdquo; (e.g. $\Delta Y$).&lt;/p&gt;
&lt;p&gt;For more details, check Theorem 1 in the paper.&lt;/p&gt;
&lt;h2 id=&#34;4-aggregation-of-mathrmattgt&#34;&gt;4. Aggregation of $\mathrm{ATT}(g,t)$&lt;/h2&gt;
&lt;p&gt;Any overall summary $\theta$ is a weighted average of the cell-specific ATTs:&lt;/p&gt;
&lt;p&gt;$$
\theta=\sum_{g=2}^T \sum_{t=g}^T w_{g, t} \operatorname{ATT}(g, t), \quad \sum_{g, t} w_{g, t}=1 .
$$&lt;/p&gt;
&lt;p&gt;Common choices:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Cohort-heterogeneity: Average effect of participating in the treatment that units in group $g$ experienced,&lt;/li&gt;
&lt;/ol&gt;
 $$
\theta_S(g)=\frac{1}{T-g+1} \sum_{t=2}^T 1\{g \leq t\} \mathrm{ATT}(g,t)
$$ 
&lt;ol start=&#34;2&#34;&gt;
&lt;li&gt;
&lt;p&gt;Calendar time heterogeneity&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Event-study / dynamic treatment effects&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;5-limitation--extension&#34;&gt;5. Limitation &amp;amp; Extension&lt;/h2&gt;
&lt;p&gt;Lee &amp;amp; Wooldridge (2023) argue that Callaway &amp;amp; Sant’Anna (2021) method is &lt;strong&gt;less efficient&lt;/strong&gt; but more resilient to functional form of covariates. The following is from their working paper:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;hellip; CS (2021) method &lt;strong&gt;uses only the period just prior to the intervention&lt;/strong&gt; in defining the control group, thereby discarding potentially useful information in earlier time periods. &lt;br&gt;&lt;/br&gt;In fact, Wooldridge (2021) shows that, under the standard “error components” structure on the error, with a homoskedastic time-constant component and homoskedastic and serially uncorrelated idiosyncratic errors, the POLS estimator is both best linear unbiased (BLUE) and asymptotically efficient. These theoretical results imply that the CS (2021) estimators are inefficient under a standard set of assumptions. The simulations in Wooldridge (2021) bear this out, showing the CS approach can be very inefficient. Balanced against the loss in precision is that the CS approach can be less biased when parallel trends are violated.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;To improve the efficiency, instead of using long differences, Lee and Wooldridge (2023) use all suitable control observations in transforming the outcome variable. Specifically, &lt;mark&gt;rather than using the single period just prior to the treatment, $Y_{g-1}$, they use the &lt;strong&gt;pre-treatment average&lt;/strong&gt;, $\frac{1}{g-1}\sum_{s = 1}^{g-1}Y_s$&lt;/mark&gt;.&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;rolling method&lt;/strong&gt; (Lee and Wooldridge, 2023) transforms panel data with staggered interventions by subtracting each unit&amp;rsquo;s average outcome across &lt;strong&gt;all pre-treatment periods&lt;/strong&gt; from their outcome in the current period of interest. This transformation, combined with &lt;strong&gt;no anticipation&lt;/strong&gt; and &lt;strong&gt;parallel trends&lt;/strong&gt; assumptions, makes the treatment assignment &lt;strong&gt;unconfounded&lt;/strong&gt; for the transformed outcome in each cohort/time cross-section. With unconfoundedness holding, we can then apply &lt;strong&gt;standard treatment effects estimators&lt;/strong&gt;, including &lt;strong&gt;doubly robust methods&lt;/strong&gt; and matching, utilizing &lt;strong&gt;all not-yet-treated units&lt;/strong&gt; as the valid control group for that specific cross-section.&lt;/p&gt;
&lt;h2 id=&#34;6-r-code-example&#34;&gt;6. R code Example&lt;/h2&gt;
&lt;p&gt;The following R code example is provided by Professor &lt;a href=&#34;https://github.com/scunning1975&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Scott Cunningham&lt;/a&gt;.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;library&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;readstata13&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;library&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;ggplot2&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;library&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;did&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# Callaway &amp;amp; Sant&amp;#39;Anna&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;castle&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;data.frame&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;read.dta13&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;s&#34;&gt;&amp;#39;https://github.com/scunning1975/mixtape/raw/master/castle.dta&amp;#39;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;castle&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;$&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;effyear&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;[is.na&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;castle&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;$&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;effyear&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;]&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;0&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# untreated units have effective year of 0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Estimating the effect on log(homicide)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;atts&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;att_gt&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;yname&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;l_homicide&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# LHS variable&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;tname&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;year&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# time variable&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;idname&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;sid&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# id variable&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;gname&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;effyear&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# first treatment period variable&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;data&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;castle&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# data&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;xformla&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;kc&#34;&gt;NULL&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# no covariates&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;c1&#34;&gt;#xformla = ~ l_police, # with covariates&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;est_method&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;dr&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# &amp;#34;dr&amp;#34; is doubly robust. &amp;#34;ipw&amp;#34; is inverse probability weighting. &amp;#34;reg&amp;#34; is regression&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;control_group&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;nevertreated&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# set the comparison group which is either &amp;#34;nevertreated&amp;#34; or &amp;#34;notyettreated&amp;#34; &lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;bstrap&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;kc&#34;&gt;TRUE&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# if TRUE compute bootstrapped SE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;biters&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1000&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# number of bootstrap iterations&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;print_details&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;kc&#34;&gt;FALSE&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# if TRUE, print detailed results&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;clustervars&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;sid&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# cluster level&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;               &lt;span class=&#34;n&#34;&gt;panel&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;kc&#34;&gt;TRUE&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;c1&#34;&gt;# whether the data is panel or repeated cross-sectional&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Aggregate ATT&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;agg_effects&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;aggte&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;atts&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;type&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;group&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;summary&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;agg_effects&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Group-time ATTs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;summary&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;atts&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Plot group-time ATTs&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;ggdid&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;atts&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Event-study&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;agg_effects_es&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;aggte&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;atts&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;type&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;dynamic&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;summary&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;agg_effects_es&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Plot event-study coefficients&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;ggdid&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;agg_effects_es&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Callaway, Brantly and Pedro H. C. Sant’Anna (2021), “Difference-in-Differences with multiple time periods,” &lt;i&gt;Journal of Econometrics&lt;/i&gt;, Themed Issue: Treatment Effect 1, 225 (2), 200–230.&lt;/p&gt;
&lt;p&gt;Lee, S. J., &amp;amp; Wooldridge, J. M. (2023). A Simple Transformation Approach to Difference-in-Differences Estimation for Panel Data (SSRN Scholarly Paper No. 4516518). Social Science Research Network. &lt;a href=&#34;https://doi.org/10.2139/ssrn.4516518&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.2139/ssrn.4516518&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Sant’Anna, Pedro H. C. and Jun Zhao (2020), “Doubly robust difference-in-differences estimators,” &lt;i&gt;Journal of Econometrics&lt;/i&gt;, 219 (1), 101–22.&lt;/p&gt;
&lt;p&gt;How does doubly robust DiD estimator works? Check this: &lt;a href=&#34;https://psantanna.com/DiD/05_Covariates.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Lecture 5: How Covariates can make your DiD More Plausible&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://bcallaway11.github.io/did/index.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;did&lt;/a&gt; R package 📦&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Intuition for Doubly Robust Estimator</title>
      <link>https://chenxing.space/blog/intuition-for-doubly-robust-estimator/</link>
      <pubDate>Sun, 18 May 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/intuition-for-doubly-robust-estimator/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;To estimate ATE, we can either use outcome regression or inverse propensity weighting (IPW). While each approach has merits, combining them offers significant advantage – double robustness. In this post, I summarize the intuition for doubly robust estimator from Professor Ding&amp;rsquo;s textbook (Ding 2024), and connects this framework to debiased machine learning (DML) through Riesz representation theory. By understanding these connections, we can gain some insights into how modern causal inference methods effectively correct for bias in treatment effect estimation.&lt;/p&gt;
&lt;h2 id=&#34;two-characterizations-of-the-ate&#34;&gt;Two characterizations of the ATE&lt;/h2&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Basic setting)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
SUTVA, unconfoundedness and overlap

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Let $Y(0), Y(1)$ be potential outcomes and $D$ be a binary treatment variable. Consider ATE.&lt;/p&gt;
&lt;p&gt;First, we can use the outcome regression,&lt;/p&gt;

$$
\tau = \E\{\mu_1(X) - \mu_2(X) \},
$$

where

$$
\mu_1 = \E\{Y(1) \mid X \} = \E\{Y \mid D = 1, X \}, 
$$

$$
\mu_0 = \E\{Y(0) \mid X \} = \E\{Y \mid D = 0, X \}
$$

&lt;p&gt;Second, we can use the inverse propensity score weighting (IPW) approach,&lt;/p&gt;

$$\tau = \E\left\{\frac{DY}{e(X)} \right\} - \E\left\{\frac{(1-D)Y}{1-e(X)} \right\},$$

&lt;p&gt;where $e(X) = \P(D = 1 \mid X)$ is the propensity score.&lt;/p&gt;
&lt;p&gt;It completely ignores the outcome model. However, if there exist covariates $X$ that are predictive of $Y$, then even when the outcome model is misspecified, including it can reduce the variance compared to using IPW alone.&lt;/p&gt;
&lt;h2 id=&#34;key-insights&#34;&gt;Key Insights&lt;/h2&gt;
&lt;p&gt;This motivates combining the two approaches to:&lt;/p&gt;
&lt;p&gt;&lt;mark&gt;1. Reduce the variance of the IPW estimator&lt;/mark&gt;&lt;/p&gt;
&lt;p&gt;&lt;mark&gt;2. Reduce the bias of the outcome regression&lt;/mark&gt;&lt;/p&gt;
&lt;h3 id=&#34;reducing-the-variance&#34;&gt;Reducing the Variance&lt;/h3&gt;

$$
\mu_1=E\{Y(1)\}=E\left\{Y(1)-\mu_1\left(X, \beta_1\right)\right\}+E\left\{\mu_1\left(X, \beta_1\right)\right\} .
$$

&lt;p&gt;&lt;strong&gt;Idea&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;View &lt;mark&gt;$Y-\mu_1\left(X, \beta_1\right)$&lt;/mark&gt; as a &amp;ldquo;pseudo potential outcome&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Then apply IPW to it:&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;

$$
\begin{aligned}
\mu_1 &amp; = E\left\{\frac{D\left\{Y-\mu_1\left(X, \beta_1\right)\right\}}{e(X)}\right\}+E\left\{\mu_1\left(X, \beta_1\right)\right\} \\
&amp; =E\left\{\frac{D\left\{Y-\mu_1\left(X, \beta_1\right)\right\}}{e(X)}+\mu_1\left(X, \beta_1\right)\right\},
\end{aligned}
$$

&lt;p&gt;Similarly,&lt;/p&gt;

$$
\begin{aligned}
\mu_0 &amp; =E\left\{\frac{(1-D)\left\{Y-\mu_0\left(X, \beta_0\right)\right\}}{1-e(X)}\right\}+E\left\{\mu_0\left(X, \beta_0\right)\right\} \\
&amp; =E\left\{\frac{(1-D)\left\{Y-\mu_0\left(X, \beta_0\right)\right\}}{1-e(X)}+\mu_0\left(X, \beta_0\right)\right\},
\end{aligned}
$$

&lt;p&gt;Notice that,&lt;/p&gt;

$$
\begin{aligned}
\mu_1 - \mu_0 &amp; = E\left\{\frac{D\left\{Y-\mu_1\left(X, \beta_1\right)\right\}}{e(X)}+\mu_1\left(X, \beta_1\right)\right\} \\ &amp;\quad - E\left\{\frac{(1-D)\left\{Y-\mu_0\left(X, \beta_0\right)\right\}}{1-e(X)}+\mu_0\left(X, \beta_0\right)\right\} 
\end{aligned}
$$

&lt;p&gt;$\mu_1 - \mu_2$ gives us the AIPW estimator.&lt;/p&gt;
&lt;h3 id=&#34;reducing-the-bias&#34;&gt;Reducing the Bias&lt;/h3&gt;

$$
\mu_1 =E\left\{\frac{Z\left\{Y-\mu_1\left(X, \beta_1\right)\right\}}{e(X)}\right\}+E\left\{\mu_1\left(X, \beta_1\right)\right\} 
$$

&lt;p&gt;&lt;strong&gt;Idea&lt;/strong&gt;: We can view &lt;mark&gt;$Y-\mu_1\left(X, \beta_1\right)$&lt;/mark&gt; as the regression residuals, from which we apply IPW to extract useful signals. Alternatively, we can view &lt;mark&gt;$\mu_1\left(X, \beta_1\right) - Y$&lt;/mark&gt; as the &lt;strong&gt;bias&lt;/strong&gt;, which we then use IPW to &lt;strong&gt;correct the bias&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&#34;connecting-to-ddml-riesz-representation-for-bias-correction&#34;&gt;Connecting to DDML: Riesz Representation for Bias Correction&lt;/h3&gt;
&lt;p&gt;In the generic debiased framework (Chernozhukov, Newey, and Singh 2022), we leverage the &lt;mark&gt;&lt;strong&gt;Riesz representer to correct for bias&lt;/strong&gt;&lt;/mark&gt;, similar to the approach described above. This method parallels our use of IPW to extract signals from residuals, but specifically employs the Riesz representer for bias correction.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Data is $Z = \{Y, D, X\}$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let $g$ be outcome regression, $g(D, X)=E[Y \mid D, X]$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Suppose moment is of the form: for some moment $m()$ that is &lt;strong&gt;linear in $g$&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;

$$\theta-E[m(Z ; g)]=0,$$
$$\theta = E[m(Z ; g)] :=E[g(1, X)-g(0, X)]$$

&lt;p&gt;Then debiased version of the moment is of the form:&lt;/p&gt;
&lt;p&gt;$$
\theta-E[m(Z ; g)+a(X) \cdot(Y-g(X))]=0,
$$&lt;/p&gt;
&lt;p&gt;where&lt;/p&gt;
&lt;p&gt;$a(X)$ is the Riesz Representer of the linear functional $L(g) := E[m(Z ; g)]$. The existence of $a(X)$ is guaranteed by the &lt;a href=&#34;https://en.wikipedia.org/wiki/Riesz_representation_theorem&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Riesz representation theorem&lt;/a&gt;, for all square integrable $g(\cdot)$, we have&lt;/p&gt;
&lt;p&gt;$$
E[m(Z ; g)]=E[a(X) \cdot g(X)].
$$&lt;/p&gt;
&lt;p&gt;As we consider the ATE, the Riesz Representer are just some &amp;ldquo;inverse propensity score terms&amp;rdquo;,&lt;/p&gt;
&lt;p&gt;$$
a(D, X):=\frac{D}{e(X)}-\frac{(1-D)}{1-e(X)},
$$&lt;/p&gt;
&lt;p&gt;Consider the expression&lt;/p&gt;
&lt;p&gt;$$
E[m(Z ; g)+a(X) \cdot(Y-g(X))]
$$&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    The key intuition here is that &lt;mark&gt;$Y−g(X)$&lt;/mark&gt; represents the residual part, and we use the Riesz Representer $a(X)$ to &lt;strong&gt;correct this bias/residual term&lt;/strong&gt;.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;double-robust-estimation-of-att&#34;&gt;Double robust estimation of ATT&lt;/h3&gt;
&lt;p&gt;How about the double robust estimator for ATT? Can we also use the same idea to derive it? Yes!&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;( &amp;#34;one-sided&amp;#34; unconfoundedness and overlap)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$$
D \indep Y(0) \mid X \text { and } e(X)&lt;1 .
$$

  &lt;/div&gt;
&lt;/div&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
Under the &#34;one-sided&#34; unconfoundedness and overlap assumption, 

$$
\E\{Y(0) \mid D=1\}=\frac{1}{e}\E\left\{\frac{e(X)}{1-e(X)} (1-D) Y\right\} \tag{1}
$$

and

$$
\tau_{\mathrm{T}}=\E(Y \mid D=1)-\E\left\{\frac{e(X)}{e} \frac{1-D}{1-e(X)} Y\right\},
$$

where $e=\P(D=1)$ is the marginal probability of the treatment.

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;We also have a doubly robust estimator for $\E\{Y(0) \mid D=1\}$  which combines the propensity score and the outcome models.&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 2&lt;/strong&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
     Define
$$
\tilde{\mu}_{0 \mathrm{T}}^{\mathrm{dr}} := \frac{1}{e} \E\left[ \frac{e(X, \alpha)}{1 - e(X, \alpha)}(1-D)\left\{Y-\mu_0\left(X, \beta_0\right)\right\}+D \mu_0\left(X, \beta_0\right)\right] ,
$$

Under above Assumption, if either $e(X, \alpha)=e(X)$ or $\mu_0\left(X, \beta_0\right)=\mu_0(X)$, then $\tilde{\mu}_{0 \mathrm{T}}^{\mathrm{dr}}= \E\{Y(0) \mid D=1\}$.

  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;How to come up with $\tilde{\mu}_{0 \mathrm{T}}^{\mathrm{dr}}$ ? Exactly the same idea as before!&lt;/p&gt;

$$
\E\{Y(0) \mid D=1\} = \E\{Y(0) - \mu_0\left(X, \beta_0\right) \mid D=1\} + \E\{\mu_0\left(X, \beta_0\right) \mid D=1\} 
$$

&lt;p&gt;Now, we can view &lt;mark&gt;$Y(0) - \mu_0\left(X, \beta_0\right)$&lt;/mark&gt; as a &amp;ldquo;pseudo potential outcome&amp;rdquo; under the control and apply eqn (1) to weight it, then we can get the form of $\tilde{\mu}_{0 \mathrm{T}}^{\mathrm{dr}}$ .&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Ding, Peng (2024) A first course in causal inference. &lt;a href=&#34;https://www.routledge.com/A-First-Course-in-Causal-Inference/Ding/p/book/9781032758626&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Chapman &amp;amp; Hall&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K. Newey, and Rahul Singh (2022), “Automatic Debiased Machine Learning of Causal and Structural Effects,” &lt;i&gt;Econometrica&lt;/i&gt;, 90 (3), 967–1027.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Adjust Censoring and Confounding Bias by IP Weighting</title>
      <link>https://chenxing.space/blog/how-to-adjust-censoring-bias-and-confounding-bias-with-ip-weights/</link>
      <pubDate>Tue, 01 Oct 2024 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/how-to-adjust-censoring-bias-and-confounding-bias-with-ip-weights/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;In the context of causal inference, adjusting for both censoring bias and confounding bias is crucial, particularly in survival analysis where right-censored data often complicates causal effect estimation. Right censoring occurs when the outcome of interest (e.g., time to an event) is not observed within the study period for some subjects, making it challenging to correctly assess the causal effect of a treatment. Moreover, confounding bias arises when treatment assignment is influenced by pre-treatment covariates, potentially leading to biased estimates of treatment effects if not properly accounted for.&lt;/p&gt;
&lt;p&gt;To address these issues, &lt;mark&gt;&lt;strong&gt;inverse probability weights (IPW)&lt;/strong&gt;&lt;/mark&gt; are commonly employed. IPW adjusts for confounding by reweighting observations based on their treatment probabilities given covariates. In addition, &lt;mark&gt;&lt;strong&gt;inverse probability of censoring weights (IPCW)&lt;/strong&gt;&lt;/mark&gt; further &lt;strong&gt;adjust for censoring by reweighting observations based on their probabilities of being uncensored&lt;/strong&gt;. Together, these techniques allow us to estimate causal effects in the presence of both confounding and censoring biases, providing more accurate insights into the relationships between treatments and outcomes.&lt;/p&gt;
&lt;p&gt;This post explains how to apply IP weights to adjust for these biases, with references to key resources.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Hernán and Robins’ &lt;em&gt;Causal Inference: What If&lt;/em&gt; (2020), especially &lt;em&gt;Chapter 8.5&lt;/em&gt; and &lt;em&gt;Chapter 12.6&lt;/em&gt;, covers censoring and missing data in detail&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Zubizarreta et al.’s &lt;em&gt;Handbook of Matching and Weighting Adjustments for Causal Inference&lt;/em&gt; (2023), Chapter 21, discusses treatment heterogeneity in survival outcomes. These resources form the basis for our explanation on using IP weights in survival analysis&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;ip-weighting&#34;&gt;IP Weighting&lt;/h2&gt;
&lt;p&gt;Imagine we want to estimate the causal effect of eco-friendly packaging on a product’s selling price. However, the selling price is right-censored — meaning we only observe the price for items that have been sold, while unsold items remain censored (i.e., their final selling price is unknown).&lt;/p&gt;
&lt;h3 id=&#34;setting&#34;&gt;Setting:&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Let $W \in \{0,1\}$ be a binary &lt;strong&gt;treatment&lt;/strong&gt; variable (e.g., $W= 1$, if the item has a eco-friendly package; $W = 0$, otherwise)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let $Y(w, s)$ be potential &lt;strong&gt;outcomes&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let $S = \textbf{1} \{C = 0\}$ be &lt;strong&gt;non-censoring indicator&lt;/strong&gt;, where $C \in \{0,1\}$ is the censoring indicator&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;e.g. $S = 1$, if the item is non-censored (in this case sold); $S = 0$, if the item is censored (in this case on-sale)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Let $X, L \in \mathcal{X}$ be two sets of &lt;strong&gt;covariates&lt;/strong&gt; and $X \subset L$. Then there exist a function $f(\cdot)$ such that $X = f(L)$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Considering $Y(W = w, S = s)$, our analysis was necessarily restricted to uncensored individuals, i.e., those with $S = 1$, because those were the only ones with known values of the outcome $Y$. Thus, the causal effect of interest is:&lt;/p&gt;

$$
\tau = \mathbb{E}\{Y(w = 1, s = 1)\} - \mathbb{E}\{Y(w = 0, s= 1)\}
$$

&lt;h3 id=&#34;assumptions&#34;&gt;Assumptions:&lt;/h3&gt;
&lt;p&gt;Recall the three identification conditions in &lt;em&gt;Chapter 8.5&lt;/em&gt; (Hernán and Robins, 2020, p. 113-114). Note that, the book uses $Y, A, C, L$ to represent the outcome, treatment, censoring indicator and covariates respectively.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;First, the average outcome in the uncensored individuals must equal the unobserved average outcome in the censored individuals with the same values of $A$ and $L$. This provision will be satisfied &amp;hellip; if the variables in $A$ and $L$ are sufficient to block all backdoor paths between $C$ and $Y$.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Second, IP weighting requires that all conditional probabilities of being uncensored given A and the variables in L must be greater than zero.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;The third condition is consistency, including sufficiently well-deﬁned interventions. IP weighting is used to create a pseudo-population in which censoring $C$ has been abolished, and in which the effect of the treatment $A$ is the same as in the original population.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Following above three identification conditions, we make the following assumptions. Denote the &lt;strong&gt;propensity score&lt;/strong&gt; as $e(\cdot)$ and &lt;strong&gt;censoring score&lt;/strong&gt; as $c(\cdot)$,&lt;/p&gt;

$$
\begin{alignat}{2}
    &amp;\text{1. (Unconfoundedness):} \quad &amp;&amp; \{Y(w= 0, s = 1), Y(w= 1, s = 1)\} \perp \!\!\! \perp W \mid X \\[1.5em]
    &amp;\text{2. (Overlap):} \quad &amp;&amp; 0 &lt; e(X) := \mathbb{P}(W = 1 \mid X) &lt; 1 \\[1.5em]
    &amp;\text{3. (Ignorable censoring):} \quad &amp;&amp; \{Y(w= 0, s = 1), Y(w= 1, s = 1)\} \perp \!\!\! \perp S \mid (W, L) \\[1.5em]
    &amp;\text{4. (Positivity):} \quad &amp;&amp; c(W, L) := \mathbb{P}(S = 1 \mid (W, L)) &gt; 0
\end{alignat}
$$

&lt;h3 id=&#34;claim&#34;&gt;Claim:&lt;/h3&gt;

$$
\begin{align}
    \tau &amp;= \mathbb{E}\left[Y(w= 1, s = 1) - Y(w= 0, s = 1)\right] \\[1em]
    &amp;= \mathbb{E}\left[\frac{S W Y}{c(W, L) e(X)} - \frac{S (1-W) Y}{c(W, L) (1-e(X))}\right] \quad \tag{2}
\end{align}
$$

&lt;h3 id=&#34;proof&#34;&gt;Proof:&lt;/h3&gt;
&lt;p&gt;First, note that the eqn (2) is well defined because of A2 and A4.&lt;/p&gt;
&lt;p&gt;I only give the proof for,
$$\mathbb{E}\left[Y(w= 1, s = 1)\right] = \mathbb{E}\left[\frac{S W Y}{c(W, L) e(X)}\right],$$ as the $\mathbb{E}\left[Y(w= 0, s = 1)\right]$ part follows the same logic.&lt;/p&gt;
&lt;p&gt;Recall that,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;By our setting, $X = f(L)$ for some known function $f$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$c(W, L) := \mathbb{P}(S = 1 \mid W, L) =  \mathbb{E}\left[S \mid W,L \right]$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Then we have,&lt;/p&gt;

\begin{aligned}
    \mathbb{E}\left[\frac{S W Y}{c(W, L) e(X)}\right] 
    &amp;= \mathbb{E}\left\{\frac{W}{e(f(L))c(W,L))} \cdot \mathbb{E}\left[SY(w, s=1) \,\middle|\, W,L \right] \right\} &amp;&amp;\text{(by LIE)} \\[1.5em]
    &amp; = \mathbb{E}\left\{\frac{W}{e(f(L))c(W,L))} \cdot \mathbb{E}\left[S \,\middle|\, W,L \right] \cdot \mathbb{E}\left[Y(w, s=1) \,\middle|\, W,L \right] \right\} &amp;&amp;\text{(by A3)} \\[1.5em]
    &amp; = \mathbb{E}\left\{\frac{W}{e(f(L))} \cdot \mathbb{E}\left[Y(w, s=1) \,\middle|\, W,L \right] \right\}  \\[1.5em]
    &amp; = \mathbb{E}\left\{\mathbb{E}\left[\frac{W}{e(f(L))} \cdot Y(w=1, s=1) \,\middle|\, W,L \right] \right\} &amp;&amp;\text{(by SUTVA)}  \\[1.5em]
    &amp; = \mathbb{E}\left[\frac{W}{e(X)} \cdot Y(w=1, s=1) \right]  \\[1.5em]
    &amp; = \mathbb{E}\left[Y(w=1, s=1) \right] &amp;&amp;\text{(by A1 and IPW)}
\end{aligned}

&lt;p&gt;The last equation holds by the classical proof of the unbiasedness of inverse propensity score weighting (IPW) estimator. For more details, one can check Theorem 11.3 at page 158-159 in Peng Ding&amp;rsquo;s textbook (Ding, 2023).&lt;/p&gt;
&lt;p align=&#39;right&#39;&gt;Q.E.D.&lt;/p&gt;
&lt;h2 id=&#34;future-work&#34;&gt;Future Work&lt;/h2&gt;
&lt;p&gt;How to create a &lt;strong&gt;doubly robust&lt;/strong&gt; estimator based on above setting? One related literature is the causal survival forest model (Cui et al., 2023), which estimates heterogeneous treatment effects in time-to-event setting and obtains doubly robustness property. But sometimes we are more interested in the downstream outcomes (e.g. selling price) after the event (e.g. being sold). This motivates us to create a new estimator that is doubly robust&amp;hellip;&lt;/p&gt;
&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;p&gt;Hernán MA, Robins JM (2020). Causal Inference: What If. Boca Raton: Chapman &amp;amp; Hall/CRC.&lt;/p&gt;
&lt;p&gt;Zubizarreta, J. R., Stuart, E. A., Small, D. S., &amp;amp; Rosenbaum, P. R. (2023). &lt;i&gt;Handbook of Matching and Weighting Adjustments for Causal Inference&lt;/i&gt;. CRC Press.&lt;/p&gt;
&lt;p&gt;Cheng, C., Li, F., Thomas, L. E., &amp;amp; Li, F. (Frank). (2022). Addressing Extreme Propensity Scores in Estimating Counterfactual Survival Functions via the Overlap Weights. &lt;i&gt;American Journal of Epidemiology&lt;/i&gt;, &lt;i&gt;191&lt;/i&gt;(6), 1140–1151. &lt;a href=&#34;https://doi.org/10.1093/aje/kwac043&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1093/aje/kwac043&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=j1lFiviKcmM&amp;amp;t=1672s&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;YouTube tutorial about IPCW: &amp;ldquo;Survival Analysis, Censoring and Time Scales&amp;rdquo;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Cui, Y., Kosorok, M. R., Sverdrup, E., Wager, S., &amp;amp; Zhu, R. (2023). Estimating heterogeneous treatment effects with right-censored data via causal survival forests. &lt;i&gt;Journal of the Royal Statistical Society Series B: Statistical Methodology&lt;/i&gt;, &lt;i&gt;85&lt;/i&gt;(2), 179–211. &lt;a href=&#34;https://doi.org/10.1093/jrsssb/qkac001&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1093/jrsssb/qkac001&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Ding, Peng. “A First Course in Causal Inference.” arXiv, October 3, 2023. &lt;a href=&#34;https://doi.org/10.48550/arXiv.2305.18793&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.48550/arXiv.2305.18793&lt;/a&gt;.&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
