<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>causal ML | Chen Xing</title>
    <link>https://chenxing.space/tag/causal-ml/</link>
      <atom:link href="https://chenxing.space/tag/causal-ml/index.xml" rel="self" type="application/rss+xml" />
    <description>causal ML</description>
    <generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Mon, 16 Feb 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://chenxing.space/media/sharing.png</url>
      <title>causal ML</title>
      <link>https://chenxing.space/tag/causal-ml/</link>
    </image>
    
    <item>
      <title>A Road Map of Nonparametric Efficiency in Causal Inference</title>
      <link>https://chenxing.space/blog/big/</link>
      <pubDate>Mon, 16 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/big/</guid>
      <description>&lt;p&gt;Causal inference promises to answer &amp;ldquo;what if&amp;rdquo; questions from observational data. But the standard approach — specify a parametric model, fit it, report the estimate — breaks down under misspecification. And misspecification is the norm, not the exception. This post draws from my notes on (Kennedy 2023). I walk through the core ideas of nonparametric efficiency theory — the framework that lets us pair flexible machine learning with rigorous inference. The central character is the &lt;mark&gt;&lt;strong&gt;Efficient Influence Function&lt;/strong&gt;&lt;/mark&gt;, a single mathematical object that tells us how to build estimators, why they work, and when we can trust their confidence intervals.&lt;/p&gt;
&lt;h2 id=&#34;1-motivation&#34;&gt;1. Motivation&lt;/h2&gt;
&lt;p&gt;First let&amp;rsquo;s keep the &lt;strong&gt;casual&lt;/strong&gt; and &lt;strong&gt;statistical&lt;/strong&gt; issues separate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Causal to Statistical:&lt;/strong&gt; The target causal parameter (e.g., ATE) can often be identified as a statistical &lt;strong&gt;functional&lt;/strong&gt; of the observed data distribution.&lt;/p&gt;
&lt;p&gt;$$\psi: \mathcal{P} \mapsto \mathbb{R}$$&lt;/p&gt;
&lt;p&gt;Once identified, the causal problem becomes a &lt;mark&gt;pure functional estimation problem&lt;/mark&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Problem with Parametrics:&lt;/strong&gt; Parametric models are likely misspecified, leading to biased estimates of the functional.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Problem with Naive Nonparametrics:&lt;/strong&gt; While we should use flexible machine learning (nonparametric) methods to avoid misspecification, simple &amp;ldquo;plug-in&amp;rdquo; estimators (fitting models and plugging predictions into the formula) fail because:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;They are generally not $\sqrt{n}$-consistent (converge too slowly).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;They do not yield valid Confidence Intervals.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;The Solution:&lt;/strong&gt; We need Efficiency Theory to understand the theoretical limit of performance and Influence Functions to construct estimators that reach that limit.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;2-efficiency-theory-lower-bounds&#34;&gt;2. Efficiency Theory (Lower Bounds)&lt;/h2&gt;
&lt;p&gt;Now let&amp;rsquo;s build the &amp;ldquo;best&amp;rdquo; estimator.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The Benchmark:&lt;/strong&gt; We want to find the &lt;mark&gt;&lt;strong&gt;best possible performance&lt;/strong&gt;&lt;/mark&gt; (lowest Mean Squared Error) any estimator can achieve.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Parametric Intuition:&lt;/strong&gt; In parametric models, the Cramer-Rao (CR) lower bound sets this benchmark: no unbiased estimator can have a variance lower than the inverse Fisher information.&lt;/p&gt;
&lt;p&gt;But now we are in the &amp;ldquo;nonparametric world&amp;rdquo;, how to exploit CR lower bound?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The &amp;ldquo;Submodel&amp;rdquo; Trick:&lt;/strong&gt; To apply CR bounds to infinite-dimensional nonparametric models:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;We define a &lt;em&gt;Parametric Submodel&lt;/em&gt;: a smooth, one-dimensional path through the complex nonparametric space that passes through the true distribution.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Logic:&lt;/em&gt; If we cannot estimate the parameter well in this simple &amp;ldquo;parametric&amp;rdquo; slice, we certainly cannot do it in the full nonparametric model.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;em&gt;Result:&lt;/em&gt; The hardest submodel defines the Local Minimax Lower Bound.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h2 id=&#34;3-the-unified-approach-central-role-of-the-eif&#34;&gt;3. The Unified Approach: Central Role of the EIF&lt;/h2&gt;
&lt;p&gt;The &lt;strong&gt;Efficient Influence Function (EIF)&lt;/strong&gt; is the &amp;ldquo;master key.&amp;rdquo; It is the derivative term in a Distributional Taylor Expansion&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a href=&#34;#fn:1&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt; (von Mises expansion):&lt;/p&gt;







&lt;figure class=&#34;blog-figure&#34;&gt;
   &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20260216113746294.png&#34; alt=&#34;image-20260216113746294&#34; style=&#34;zoom:50%;&#34; /&gt; 
  &lt;figcaption&gt;
    
      Figure 1: Distributional Taylor Expansion
    
  &lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This single expansion solves all three statistical tasks:&lt;/p&gt;
&lt;h3 id=&#34;a-how-to-construct-the-estimator-the-recipe&#34;&gt;A. How to construct the estimator? (The Recipe)&lt;/h3&gt;
&lt;p&gt;The expansion reveals that the bias of a naive plug-in estimator is roughly the expected value of the EIF.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Recipe:&lt;/em&gt; We construct a One-Step Estimator by taking the naive plug-in estimator and adding the empirical average of the estimated EIF to &amp;ldquo;de-bias&amp;rdquo; it.&lt;/p&gt;
&lt;p&gt;$$
\hat{\psi}_{\text{one-step}} = \text{Plug-in} + \frac{1}{n}\sum \text{EIF}(\text{Data})
$$&lt;/p&gt;
&lt;details&gt;
&lt;summary&gt;Click to expand: Deconstructing the One-Step Estimator&lt;/summary&gt;
&lt;h4 id=&#34;the-intuition-a-newton-raphson-correction&#34;&gt;The Intuition: A &amp;ldquo;Newton-Raphson&amp;rdquo; Correction&lt;/h4&gt;
&lt;p&gt;Think of it like Newton&amp;rsquo;s method in calculus: if you want to find the root of a function, you make an initial guess, calculate the derivative (slope) at that point, and use it to take a &amp;ldquo;step&amp;rdquo; closer to the true answer.&lt;/p&gt;
&lt;p&gt;In this context:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The &lt;em&gt;Plug-in&lt;/em&gt; is your initial &amp;ldquo;guess&amp;rdquo; at the parameter.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;em&gt;EIF&lt;/em&gt; acts as the &amp;ldquo;derivative&amp;rdquo; (gradient) that tells you which direction to move to fix the error.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;em&gt;Formula&lt;/em&gt; represents taking that single &amp;ldquo;step&amp;rdquo; to correct the bias of your initial guess.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id=&#34;deconstructing-the-formula&#34;&gt;Deconstructing the Formula&lt;/h4&gt;
 
$$
\hat{\psi}_{\text{one-step}} = \underbrace{\psi(\hat{P})}_{\text{Plug-in}} + \underbrace{\frac{1}{n}\sum_{i=1}^n \phi(Z_i; \hat{P})}_{\text{Bias Correction}}
$$ 
&lt;p&gt;&lt;strong&gt;Part A: The Naive Plug-in&lt;/strong&gt; $\psi(\hat{P})$&lt;/p&gt;
&lt;p&gt;This is the estimate you get if you train your machine learning models (like regression or propensity scores), estimating the whole distribution $\hat{P}$, and simply &amp;ldquo;plug them in&amp;rdquo; to the formula for your parameter.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Why it fails:&lt;/em&gt; Flexible machine learning models trade bias for variance (regularization). They &amp;ldquo;smooth&amp;rdquo; the data to avoid overfitting. While this is good for predicting individual outcomes, it creates a first-order bias in the target parameter that does not vanish fast enough (slower than $1/\sqrt{n}$). If you stop here, your confidence intervals will be wrong.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Part B: The Bias Correction&lt;/strong&gt; $\frac{1}{n}\sum \text{EIF}$&lt;/p&gt;
&lt;p&gt;This term estimates the bias of the plug-in and removes it. The logic relies on the Distributional Taylor Expansion, which tells us that the error of the plug-in estimator is approximately equal to the negative expectation of the Influence Function:&lt;/p&gt;
&lt;p&gt;$$
\psi(\hat{P}) - \psi(P_{\text{true}}) \approx -\mathbb{E}[\text{EIF}]
$$&lt;/p&gt;
&lt;p&gt;Since the error is roughly $-\mathbb{E}[\text{EIF}]$, we can &amp;ldquo;cancel out&amp;rdquo; this error by adding the empirical average of the EIF estimated from our data. Note: if our initial estimate $\hat{P}$ were perfect (the truth), the average of the EIF would be exactly zero. The fact that this term is &lt;em&gt;not&lt;/em&gt; zero reflects the bias in our initial model that needs to be corrected.&lt;/p&gt;
&lt;h4 id=&#34;concrete-example-average-treatment-effect-ate&#34;&gt;Concrete Example: Average Treatment Effect (ATE)&lt;/h4&gt;
&lt;p&gt;To make this concrete, consider the ATE. This specific &amp;ldquo;One-Step&amp;rdquo; estimator is mathematically identical to the famous AIPW (Augmented Inverse Probability Weighting) or Doubly Robust estimator.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The &lt;em&gt;Plug-in:&lt;/em&gt; You train a regression model to predict outcomes for treated vs. untreated groups. You calculate the average difference. The problem? If your regression is slightly wrong (which it always is), your effect estimate is biased.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;em&gt;EIF Correction:&lt;/em&gt; You look at the residuals — the difference between what your model predicted and what actually happened, weighted by the propensity score (the probability of treatment). If your model consistently under-predicts outcomes for the treated group, the EIF term will be positive. Adding this average EIF to the plug-in &amp;ldquo;bumps&amp;rdquo; the estimate up, correcting the bias.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id=&#34;summary&#34;&gt;Summary&lt;/h4&gt;
&lt;p&gt;The &amp;ldquo;One-Step Estimator&amp;rdquo; recipe acknowledges that modern Machine Learning is great at learning patterns (the Plug-in) but bad at getting the total volume/area right (Bias). By adding the average of the Efficient Influence Function (the derivative), you explicitly calculate that missing &amp;ldquo;volume&amp;rdquo; and add it back in, ensuring the final estimate is unbiased and efficient.&lt;/p&gt;
&lt;/details&gt;
&lt;h3 id=&#34;b-how-to-analyze-the-estimator-the-proof&#34;&gt;B. How to analyze the estimator? (The Proof)&lt;/h3&gt;
&lt;p&gt;The expansion includes a &lt;em&gt;Remainder Term&lt;/em&gt; ($R_2$). To prove the estimator is $\sqrt{n}$-consistent, we must show this remainder is negligible (second-order).&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Sample Splitting:&lt;/em&gt; We need to use cross-fitting to handle the &amp;ldquo;empirical process term&amp;rdquo; (preventing overfitting when estimating nuisance parameters).&lt;/p&gt;
&lt;h3 id=&#34;c-how-to-correct-the-bias-double-robustness&#34;&gt;C. How to correct the bias? (Double Robustness)&lt;/h3&gt;
&lt;p&gt;The math of the EIF reveals that the error depends on the &lt;em&gt;product of errors&lt;/em&gt; of the nuisance parameters (e.g., error in outcome regression $\times$ error in propensity score).&lt;/p&gt;
&lt;p&gt;This yields &lt;strong&gt;Double Robustness&lt;/strong&gt;: even if individual machine learning models converge slowly (e.g., $n^{-1/4}$), their product converges fast enough ($n^{-1/2}$) to allow for valid inference.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;4-the-link-between-eif-and-riesz-representer&#34;&gt;4. The Link Between EIF and Riesz Representer&lt;/h2&gt;
&lt;p&gt;The connection between the Efficient Influence Function (EIF) and the Riesz Representer (RR) is the bridge between &lt;em&gt;theoretical statistics&lt;/em&gt; and &lt;em&gt;practical machine learning&lt;/em&gt;.&lt;/p&gt;
&lt;h3 id=&#34;the-deconstruction-of-the-eif&#34;&gt;The Deconstruction of the EIF&lt;/h3&gt;
&lt;p&gt;We already know the EIF is the &amp;ldquo;correction term&amp;rdquo; we add to a naive plug-in estimator to fix bias. The Riesz Representer is the key component &lt;em&gt;inside&lt;/em&gt; that EIF.&lt;/p&gt;
&lt;p&gt;For a vast class of problems (linear functionals), the EIF always takes this specific structure:&lt;/p&gt;
 $$
\underbrace{\phi(Z)}_{\text{EIF}} = \underbrace{\psi(P)}_{\text{Plug-in}} + \underbrace{\alpha(X, W)}_{\text{Riesz Representer}} \times \underbrace{(Y - \mu(X, W))}_{\text{Outcome Residual}} 
$$ 
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$\mu(X, W)$: The conditional expectation (the outcome model, e.g., regression of $Y$ on $X, W$).&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\alpha(X, W)$: The Riesz Representer. It tells us &lt;em&gt;how much weight&lt;/em&gt; to give each residual to correct the bias.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The EIF says: &amp;ldquo;To fix bias, look at your regression errors ($Y - \mu$).&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The RR says: &amp;ldquo;Here is exactly how to &lt;em&gt;weight&lt;/em&gt; those errors for this specific causal problem.&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;takeaway&#34;&gt;Takeaway&lt;/h2&gt;
&lt;p&gt;Here is the takeaway. Nonparametric efficiency theory is not an abstraction separate from practice — it is the practice. The EIF tells us what to estimate, the Riesz Representer tells us how to weight the errors, and double robustness explains why the whole thing holds together even when our models are imperfect. If we use AIPW, DML, or any doubly robust estimator, we are already using these ideas. The theory simply makes explicit what those estimators are doing and why they deserve our trust.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Kennedy, Edward H. (2023), “Semiparametric doubly robust targeted double machine learning: a review.”&lt;/p&gt;
&lt;p&gt;Williams, Nicholas T., Oliver J. Hines, and Kara E. Rudolph (2026), “Riesz representers for the rest of us.”&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K. Newey, Victor Quintas-Martinez, and Vasilis Syrgkanis (2024), “Automatic debiased machine learning via riesz regression.”&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K. Newey, and Rahul Singh (2022), “Automatic Debiased Machine Learning of Causal and Structural Effects,” &lt;i&gt;Econometrica&lt;/i&gt;, 90 (3), 967–1027.&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;The Distributional Taylor Expansion is essentially the standard Taylor expansion applied at the distribution scale rather than the real-number scale. It is exactly equivalent to the concept of pathwise differentiability — both formalize the idea of taking a derivative of a statistical functional with respect to perturbations of the underlying distribution.&amp;#160;&lt;a href=&#34;#fnref:1&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Calibrate Your Nuisances: A Simple Fix for Doubly Robust Inference</title>
      <link>https://chenxing.space/blog/calibrate-your-nuisances-a-simple-fix-for-doubly-robust-inference/</link>
      <pubDate>Fri, 09 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/calibrate-your-nuisances-a-simple-fix-for-doubly-robust-inference/</guid>
      <description>&lt;h2 id=&#34;doubly-robust-inference-via-calibration&#34;&gt;Doubly Robust Inference via Calibration&lt;/h2&gt;
&lt;h3 id=&#34;tldr&#34;&gt;TL;DR&lt;/h3&gt;
&lt;p&gt;Van der Laan, Luedtke, and Carone introduce &lt;strong&gt;&amp;ldquo;calibrated debiased machine learning&amp;rdquo; (calibrated DML)&lt;/strong&gt;, a method that achieves doubly robust asymptotic normality for causal inference estimators by simply adding an isotonic regression calibration step to standard DML pipelines. The key innovation is that valid inference requires only one of two nuisance functions (outcome regression or propensity score) to be estimated well—the other can converge arbitrarily slowly or even inconsistently. This bridges a long-standing gap where consistency was doubly robust but inference was not, and the method can be implemented by adding just a few lines of code to existing workflows.&lt;/p&gt;
&lt;h3 id=&#34;what-is-this-paper-about&#34;&gt;What is this paper about?&lt;/h3&gt;
&lt;p&gt;Doubly robust estimators like AIPW are popular for estimating average treatment effects because they remain consistent if either the outcome regression or propensity score is correctly specified. However, there&amp;rsquo;s a crucial asymmetry: &lt;mark&gt;while consistency requires only one nuisance to be correct, valid inference (asymptotic normality, correct confidence intervals) typically requires both nuisances to converge at sufficiently fast rates $(\approx n^{-1/4})$. When one nuisance is estimated poorly—whether due to model misspecification or slow convergence—standard confidence intervals can have incorrect coverage.&lt;/mark&gt; Prior solutions like the DR-TMLE framework required computationally intensive iterative procedures and case-by-case derivations for each new parameter.&lt;/p&gt;
&lt;h4 id=&#34;motivation&#34;&gt;Motivation&lt;/h4&gt;
&lt;p&gt;The authors ask: can we achieve doubly robust inference using a simple, general-purpose procedure that works with any machine learning estimator?&lt;/p&gt;
&lt;h3 id=&#34;what-do-the-authors-do&#34;&gt;What do the authors do?&lt;/h3&gt;
&lt;p&gt;The authors develop calibrated DML, a two-step procedure that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;step 1: takes cross-fitted nuisance estimators from any machine learning algorithm&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;step 2: calibrates them using isotonic regression before constructing the debiased estimator&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The calibration step ensures that nuisance estimates satisfy certain &lt;strong&gt;empirical orthogonality conditions&lt;/strong&gt; that linearize the bias term in the doubly robust expansion. Specifically, they calibrate the outcome regression using squared error loss and the Riesz representer (inverse propensity weights for ATE) using a tailored Riesz loss.&lt;/p&gt;
&lt;p&gt;&lt;mark&gt;The paper proves that this calibration is sufficient to achieve &amp;ldquo;doubly robust asymptotic linearity&amp;rdquo; (DRAL)—meaning the estimator is asymptotically normal whenever at least one nuisance converges at $n^{-1/4}$ rate, even if the other is inconsistent.&lt;/mark&gt; They also develop bootstrap-assisted confidence intervals that avoid estimating additional nuisance functions. Empirically, they evaluate on simulated data with deliberately misspecified nuisances and on semi-synthetic benchmarks (ACIC, IHDP, Twins, LaLonde), comparing against standard AIPW and the iterative DR-TMLE.&lt;/p&gt;
&lt;h3 id=&#34;why-is-this-important&#34;&gt;Why is this important?&lt;/h3&gt;
&lt;p&gt;Applied researchers using flexible ML methods (random forests, neural networks, gradient boosting) for nuisance estimation face a dilemma: these methods provide consistency but may not converge fast enough to guarantee valid inference, especially in moderate dimensions. &lt;mark&gt;The standard &amp;ldquo;product rate&amp;rdquo; condition requiring both nuisances to converge at $n^{-1/4}$ can fail when even one nuisance is complex or misspecified. Calibrated DML transforms this multiplicative requirement into a requirement on just the better-estimated nuisance.&lt;/mark&gt;&lt;/p&gt;
&lt;p&gt;The practical benefits are substantial: better finite-sample coverage (e.g., improving from 32% to 90% coverage in ACIC-2017 simulations), reduced bias, and a method that integrates into existing pipelines with minimal code changes.&lt;/p&gt;
&lt;p&gt;The theoretical contribution is equally significant—it establishes a novel connection between prediction calibration and causal inference validity, showing that calibration of nuisances provides the &amp;ldquo;debiasing&amp;rdquo; needed for doubly robust inference.&lt;/p&gt;
&lt;h3 id=&#34;who-should-care&#34;&gt;Who should care?&lt;/h3&gt;
&lt;p&gt;Applied researchers estimating treatment effects with observational data using ML for nuisance estimation will find immediate practical value—this method provides insurance against nuisance misspecification without computational overhead.&lt;/p&gt;
&lt;p&gt;Econometricians and biostatisticians working on semiparametric inference will appreciate the theoretical framework connecting calibration to doubly robust properties.&lt;/p&gt;
&lt;p&gt;Methodologists developing new estimands can use this as a general recipe: the approach applies to any linear functional of the outcome regression (counterfactual means, ATE, partial covariance, survival outcomes under missingness), not just the ATE.&lt;/p&gt;
&lt;p&gt;Researchers in policy evaluation, epidemiology, marketing, and tech who regularly use AIPW-style estimators should consider adopting calibrated DML as a default given its robustness benefits.&lt;/p&gt;
&lt;h3 id=&#34;do-we-have-code&#34;&gt;Do we have code?&lt;/h3&gt;
&lt;p&gt;Yes, the authors provide both R and Python implementations.&lt;/p&gt;
&lt;p&gt;An R package &lt;code&gt;calibratedDML&lt;/code&gt; and Python code are available on GitHub at &lt;a href=&#34;https://github.com/Larsvanderlaan/calibratedDML&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://github.com/Larsvanderlaan/calibratedDML&lt;/a&gt;. The paper includes complete code listings for calibrating inverse propensity weights and outcome regressions using &lt;code&gt;xgboost&lt;/code&gt; with monotonicity constraints.&lt;/p&gt;
&lt;p&gt;The implementation is straightforward: isotonic regression is performed using gradient-boosted trees with &lt;code&gt;monotone_constraints=1&lt;/code&gt; and a single boosting round, making it computationally efficient and easy to integrate into existing DML workflows.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;In summary, this paper provides a remarkably practical solution to a longstanding theoretical problem. The insight that isotonic calibration—a standard tool from prediction—can unlock doubly robust inference is both elegant and immediately actionable. For anyone running AIPW or related estimators with ML nuisances, calibrating cross-fitted estimates before debiasing is now the obvious default: it costs almost nothing computationally and provides genuine protection against the scenario where one of your nuisance models is less reliable than you hoped.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;van der Laan, Lars, Alex Luedtke, and Marco Carone (2024), “Doubly robust inference via calibration,” arXiv preprint arXiv:2411.02771.&lt;/p&gt;
&lt;p&gt;An R package &lt;code&gt;calibratedDML&lt;/code&gt; and Python code are available on GitHub at &lt;a href=&#34;https://github.com/Larsvanderlaan/calibratedDML&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://github.com/Larsvanderlaan/calibratedDML&lt;/a&gt;.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>AIPW vs. Residual-on-Residual regression: Non-Parametric Flexibility or Efficiency?</title>
      <link>https://chenxing.space/blog/aipw-vs-residual-on-residual-regression-non-parametric-flexibility-or-efficiency/</link>
      <pubDate>Wed, 11 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/aipw-vs-residual-on-residual-regression-non-parametric-flexibility-or-efficiency/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Two powerful tools in causal inference are the &lt;strong&gt;Augmented Inverse Propensity Weighting (AIPW)&lt;/strong&gt; estimator and the &lt;strong&gt;Residual-on-Residual regression&lt;/strong&gt; estimator for partially linear models. Drawing from Wager’s notes (2024), this post breaks down how these estimators work, compares their strengths and weaknesses, and offers tips for when to use each.&lt;/p&gt;
&lt;h2 id=&#34;residual-on-residual-regression&#34;&gt;Residual-on-Residual regression&lt;/h2&gt;
&lt;p&gt;The &lt;strong&gt;semiparamteric partially linear model&lt;/strong&gt; (PLM) assumes the outcome $ Y $ can be written as:&lt;/p&gt;
&lt;p&gt;$$ Y = \theta D + g(X) + \epsilon \tag{1}$$&lt;/p&gt;
&lt;p&gt;Here, $ D $ is a binary treatment, $ \theta $ is the causal parameter of interest, $ g(X) $ is some unknown function of covariates $ X $, and $ \epsilon $ is random noise with $\E[\epsilon \mid D, X] = 0$. The residual-on-residual estimator (or Robinson (1988) estimator) isolates $ \theta $ in two steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Partialing out covariates&lt;/strong&gt;: &amp;ldquo;Partialing out&amp;rdquo; $ X $ from $D$ and $Y$ using nonparametric regression and cross-fitting: $ \tilde{D} = D - \hat{\E}[D \mid X] $ and $ \tilde{Y} = Y - \hat{\E}[Y \mid X] $.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Estimate the effect&lt;/strong&gt;: Use linear regression $ \tilde{Y} \sim \tilde{D} $ to get $ \hat{\theta}$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    However, the model is &lt;strong&gt;not fully general&lt;/strong&gt;, because it  imposes a &lt;strong&gt;parametric specification&lt;/strong&gt; on the key component of interest. It imposes &lt;strong&gt;additivity&lt;/strong&gt; in $g(X)$ and $D$.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;augmented-inverse-propensity-weighting-aipw-estimator&#34;&gt;Augmented Inverse Propensity Weighting (AIPW) Estimator&lt;/h3&gt;
&lt;p&gt;AIPW takes a &lt;strong&gt;fully non-parametric&lt;/strong&gt; approach, aiming to estimate ATE, $ \tau = E[Y(1) - Y(0)] $, where $ Y(1) $ and $ Y(0) $ are potential outcomes under treatment and control.&lt;/p&gt;
 $$ \hat{\tau}_{\text{AIPW}} = \frac{1}{n} \sum_{i=1}^n \left[ \hat{\mu}_1(X_i) - \hat{\mu}_0(X_i) + D_i \frac{Y_i - \hat{\mu}_1(X_i)}{\hat{e}(X_i)} - (1 - D_i) \frac{Y_i - \hat{\mu}_0(X_i)}{1 - \hat{e}(X_i)} \right] $$ 
&lt;p&gt;Here, $ \hat{\mu}_1(X) $ and $ \hat{\mu}_0(X) $ are outcome regression estimators for treated and untreated units, and $ \hat{e}(X) $ is the estimator of propensity score. AIPW combines outcome modeling with inverse propensity weighting, making it &lt;strong&gt;doubly robust&lt;/strong&gt;: it’s consistent if &lt;em&gt;either&lt;/em&gt; the outcome model or propensity score is consistent.&lt;/p&gt;
&lt;h2 id=&#34;the-key-difference-non-parametric-vs-partially-linear&#34;&gt;The Key Difference: Non-Parametric vs. Partially Linear&lt;/h2&gt;
&lt;p&gt;&lt;mark&gt;&lt;strong&gt;AIPW is fully non-parametric&lt;/strong&gt;, imposing no specific parametric form on the treatment effect, while &lt;strong&gt;residual-on-residual regression estimator assumes a partially linear structure&lt;/strong&gt;&lt;/mark&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;AIPW’s flexibility&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;AIPW does not assume a specific parametric form!&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;AIPW is &lt;strong&gt;efficient&lt;/strong&gt; in the generic &lt;strong&gt;non-parametric&lt;/strong&gt; setting.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250611092857556.png&#34; alt=&#34;image-20250611092857556&#34; style=&#34;zoom:80%;&#34; /&gt;
&lt;/blockquote&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Residual-on-residual regression’s structure&lt;/strong&gt;: What partially linear assumption buys us is that residual-on-residual estimators that exploit this constraint can have &lt;strong&gt;smaller variance than AIPW&lt;/strong&gt;. In other words, adding this additional structure makes the residual-on-residual estimator more efficient than AIPW.&lt;/p&gt;
&lt;blockquote&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250611092145259.png&#34; alt=&#34;image-20250611092145259&#34; style=&#34;zoom:80%;&#34; /&gt;
&lt;/blockquote&gt;
&lt;div class=&#34;alert alert-warning&#34;&gt;
&lt;div&gt;
  A risk   of using the residual-on-residual estimator is that &lt;strong&gt;constant treatment effect&lt;/strong&gt; model (1) may be misspecified.
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;
&lt;h2 id=&#34;why-this-matters&#34;&gt;Why This Matters&lt;/h2&gt;
&lt;p&gt;The choice between AIPW and residual-on-residual regression reflects a deeper trade-off in causal inference: &lt;strong&gt;flexibility versus efficiency&lt;/strong&gt;. AIPW’s non-parametric nature makes it a Swiss Army knife for complex data, while PLM structure is like a precision tool—effective when conditions are right.&lt;/p&gt;
&lt;p&gt;As Wager’s notes highlight:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Both AIPW and residual-on-residual regression are &lt;strong&gt;Neyman-orthogonal&lt;/strong&gt;, making them robust to first‐stage errors.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;However, their assumptions shape their performance. AIPW attains the lowest possible asymptotic variance for ATE under unconfoundedness. The residual-on-residual estimator, by imposing extra structure, can go beyond that bound when its structure is correct but at the cost of vulnerability to misspecification.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;conclusion&#34;&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;AIPW’s fully non-parametric approach offers robustness and flexibility, while residual-on-residual regression’s partially linear structure prioritizes efficiency when assumptions hold.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Wager, S. (2024). &lt;em&gt;Causal inference: A statistical learning approach&lt;/em&gt;. &lt;a href=&#34;https://web.stanford.edu/~swager/causal_inf_book.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://web.stanford.edu/~swager/causal_inf_book.pdf&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>What is the Critical Radius in Riesz Regression?</title>
      <link>https://chenxing.space/blog/what-is-the-critical-radius-in-riesz-regression/</link>
      <pubDate>Sat, 07 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/what-is-the-critical-radius-in-riesz-regression/</guid>
      <description>&lt;h2 id=&#34;tldr&#34;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;In the paper &lt;em&gt;Automatic Debiased Machine Learning via Riesz Regression&lt;/em&gt; by Chernozhukov et al. (2024), the &lt;strong&gt;critical radius&lt;/strong&gt; is a key concept for understanding how accurately machine learning (ML) can estimate the Riesz representer—a function used to debias ML estimates of economic parameters like treatment effects.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;In Simple Terms&lt;/strong&gt;: &lt;mark&gt;The critical radius measures how &amp;ldquo;hard&amp;rdquo; it is to learn a function (like the Riesz representer) from data.&lt;/mark&gt;&lt;/p&gt;
&lt;p&gt;The critical radius depends on:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Model Complexity&lt;/strong&gt;: More flexible ML models (e.g., deep neural nets) have a larger critical radius, meaning they need more data to estimate accurately.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sample Size&lt;/strong&gt;: More data reduces the critical radius, making estimates more precise.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Intuition&lt;/strong&gt;: &lt;mark&gt;Think of the critical radius as a speed limit on how fast your machine learning estimator can converge to the true function.&lt;/mark&gt; If you’re estimating a complicated function (like a highly nonlinear Riesz representer) with a flexible model (e.g., a deep neural net), the critical radius will be larger, meaning you need more data to get a good estimate. If the function is simpler or you use a less flexible model (e.g., a linear model), the critical radius is smaller, and you need less data.&lt;/p&gt;
&lt;p&gt;In the paper, the &lt;strong&gt;critical radius&lt;/strong&gt; is critical because it ensures the Riesz regression estimator is reliable enough to debias the final parameter estimate (e.g., ATE). If the critical radius is too large (due to an overly complex model or too little data), the error in estimating $\alpha_0$ could undermine the debiasing process, leading to biased or imprecise results.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why It Matters&lt;/strong&gt;: The critical radius helps bound the error in estimating the Riesz representer (see Theorem 2.1). A smaller radius means lower error, ensuring reliable debiasing of ML estimates. For example, in neural nets, the radius scales with network size and shrinks with more data. For example, $$\delta_n \propto \sqrt{\frac{\text{network size}}{n}}$$&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Practical Takeaway&lt;/strong&gt;: Choose ML models with complexity suited to your sample size to keep the critical radius small, ensuring accurate estimates. The paper’s simulations show better performance with larger samples (e.g., $n=10,000$), where the critical radius is smaller, leading to robust results.&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;critical radius&lt;/strong&gt; in the context of the paper &lt;em&gt;Automatic Debiased Machine Learning via Riesz Regression&lt;/em&gt; by Chernozhukov et al. is a statistical concept from learning theory, but its name does indeed suggest a geometric intuition. Below, I’ll explain whether the critical radius has a geometric meaning, why it’s called a &amp;ldquo;radius,&amp;rdquo; and how this relates to the paper, keeping the explanation concise and accessible for a PhD-level economist familiar with econometrics but not necessarily advanced statistical learning theory.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;does-the-critical-radius-have-a-geometric-meaning&#34;&gt;Does the Critical Radius Have a Geometric Meaning?&lt;/h2&gt;
&lt;p&gt;The critical radius has a &lt;strong&gt;geometric interpretation&lt;/strong&gt;, though it’s rooted in the abstract geometry of function spaces rather than physical space. In statistical learning, the critical radius is related to the &lt;strong&gt;complexity&lt;/strong&gt; of a set of functions (e.g., the class of neural nets or random forests used to estimate the Riesz representer). Geometrically, you can think of it as a measure of the &amp;ldquo;size&amp;rdquo; or &amp;ldquo;spread&amp;rdquo; of this function class in a high-dimensional space, which determines how hard it is to learn a specific function (like the Riesz representer $\alpha_0$) from data.&lt;/p&gt;
&lt;p&gt;Here’s the intuition:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Imagine the set of possible functions $\mathcal{A}_n$ (e.g., all possible neural nets with a given architecture) as a &amp;ldquo;cloud&amp;rdquo; of points in a function space, where each point is a function $\alpha(x)$.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The &lt;strong&gt;critical radius&lt;/strong&gt; ($\delta_n$) is like the &lt;strong&gt;radius of a ball&lt;/strong&gt; around the true function ($\alpha_0$) that captures how spread out or complex the function class is. A larger radius means the function class is more complex (e.g., deeper neural nets with many parameters), making it harder to pinpoint the true function with limited data.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;This &amp;ldquo;ball&amp;rdquo; isn’t in physical space but in a mathematical space where distance is measured by the mean square error. For example, $$|\alpha - \alpha_0|^2 = \mathbb{E}[(\alpha(X) - \alpha_0(X))^2]$$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In the paper, the critical radius is used in &lt;strong&gt;Theorem 2.1&lt;/strong&gt; to bound the error of the Riesz regression estimator:&lt;/p&gt;
&lt;p&gt;$$
|\hat{\alpha} - \alpha_0|^2 \leq C \left( M \delta_n^2 + |\alpha^* - \alpha_0|^2 + \frac{M \ln(1/\zeta)}{n} \right),
$$&lt;/p&gt;
&lt;p&gt;where $\delta_n$ is the critical radius of the function class $\mathcal{A}_n$. Geometrically, $\delta_n$ quantifies the &amp;ldquo;width&amp;rdquo; of the function class, affecting how close the estimated function $\hat{\alpha}$ can get to the true $\alpha_0$.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;why-is-it-called-a-radius&#34;&gt;Why Is It Called a &amp;ldquo;Radius&amp;rdquo;?&lt;/h2&gt;
&lt;p&gt;The term &amp;ldquo;radius&amp;rdquo; comes from its connection to &lt;strong&gt;Rademacher complexity&lt;/strong&gt; or similar complexity measures in statistical learning theory, which often have a geometric flavor. Here’s why:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Rademacher Complexity&lt;/strong&gt;: The critical radius is closely tied to the Rademacher complexity of a function class, which measures the ability of the class to fit random noise. It’s like asking, &amp;ldquo;How big is the set of functions in terms of their ability to wiggle around and fit data?&amp;rdquo; This is often visualized as the radius of a ball in a function space that encloses the class’s variability.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Covering Numbers&lt;/strong&gt;: Another related concept is the covering number, which counts how many small balls (of radius $\delta$) are needed to cover the function class. The critical radius is the smallest $\delta_n$ where the function class’s complexity balances with the sample size $n$, resembling the radius of these covering balls.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Historical Naming&lt;/strong&gt;: The term &amp;ldquo;radius&amp;rdquo; is borrowed from empirical process theory (e.g., Foster and Syrgkanis, 2019, cited in the paper), where it describes the scale of stochastic fluctuations in a function class. It’s called a radius because it behaves like the size of a region in function space where the estimator is likely to lie.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In the paper, for a neural net with width $K$ and depth $m$, the critical radius is bounded as:
$$
\delta_n \leq C \sqrt{\frac{K^2 m^2 \ln(K^2 m) \ln(n)}{n}}.
$$
This reflects the geometric idea that a more complex model (larger $K$ or $m$) has a larger &amp;ldquo;radius,&amp;rdquo; requiring more data ($n$) to shrink the error.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;summary&#34;&gt;Summary&lt;/h2&gt;
&lt;p&gt;The &lt;strong&gt;critical radius&lt;/strong&gt; has a geometric meaning as the &amp;ldquo;size&amp;rdquo; of a function class in an abstract space, reflecting its complexity and the data needed to learn a function accurately. It’s called a &amp;ldquo;radius&amp;rdquo; due to its connection to Rademacher complexity and covering numbers, which describe the spread of functions like a ball’s radius. In the paper, it quantifies the error in estimating the Riesz representer, ensuring effective debiasing of ML-based economic parameters.&lt;/p&gt;
&lt;/xaiArtifact&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;p&gt;Chernozhukov, V., Newey, W. K., Quintas-Martinez, V., &amp;amp; Syrgkanis, V. (2024). Automatic debiased machine learning via riesz regression. &lt;a href=&#34;https://arxiv.org/abs/2104.14737&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://arxiv.org/abs/2104.14737&lt;/a&gt;&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Notes on DML for DiD: A Unified Approach</title>
      <link>https://chenxing.space/blog/notes-on-dml-for-did/</link>
      <pubDate>Mon, 02 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-on-dml-for-did/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;This blog post explores how Double Machine Learning (DML) extends to conditional Difference-in-Differences (DiD), focusing on doubly robust estimators. The &lt;strong&gt;key insight&lt;/strong&gt; is that conditional DiD can be understood through the lens of cross-sectional ATT estimation.&lt;/p&gt;
&lt;h2 id=&#34;foundation-cross-sectional-att-estimation&#34;&gt;Foundation: Cross-Sectional ATT Estimation&lt;/h2&gt;
&lt;p&gt;To build intuition, we start with the familiar cross-sectional setting. Standard identification requires three assumptions: &lt;strong&gt;SUTVA, unconfoundedness, and overlap&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&#34;step-1-propensity-score-approach-for-att&#34;&gt;Step 1: Propensity Score Approach for ATT&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Unlike ATE, ATT estimation requires only &amp;ldquo;one-sided&amp;rdquo; unconfoundedness and overlap conditions.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Identification Assumptions)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
$$Z \indep Y(0) \mid X \text{ and } e(X) &lt; 1$$

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Estimate ATT using &lt;strong&gt;IPW&lt;/strong&gt;&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Ding (2024), Section 13.2)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602144838745.png&#34; alt=&#34;image-20250602144838745&#34; style=&#34;zoom:40%;&#34; /&gt; 

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;More general, Li et al. (2018a) gave a unified discussion of the causal estimands in observational studies.&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(Ding (2024), Section 13.4)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602144337279.png&#34; alt=&#34;image-20250602144337279&#34; style=&#34;zoom:40%;&#34; /&gt;
&lt;br&gt;
Summary Table of common estimands: 
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602145602468.png&#34; alt=&#34;image-20250602145602468&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;This table provides us a good way to understand and remember IPW estimator for ATT&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How to remember $\tau^h$? Apply IPW on &lt;mark&gt;&amp;ldquo;pseudo outcome&amp;rdquo; $Yh(X)$ &lt;/mark&gt; then divide by $E(h(X))$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;When the parameter of interest is ATT, then $$E(h(X)) = E(e(X)) = E(E(Z \mid X)) = E(Z) = \P(Z = 1) = e$$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Use it to better understand IPW for ATT&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;step-2-doubly-robust-att-estimator&#34;&gt;Step 2: Doubly Robust ATT Estimator&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Combines outcome regression and IPW methods&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;For DR estimator of ATT, check &lt;a href=&#34;https://chenxing.space/blog/intuition-for-doubly-robust-estimator/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;my previous post&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;More generally, we have&lt;/p&gt;







&lt;div class=&#34;math-environment theorem&#34; id=&#34;theorem-3&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Theorem 3&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(DR for general estimand, see Ding (2024), page 191)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602151637808.png&#34; alt=&#34;image-20250602151637808&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;extension-to-conditional-did&#34;&gt;Extension to Conditional DiD&lt;/h2&gt;
&lt;h3 id=&#34;identification-assumptions&#34;&gt;Identification Assumptions&lt;/h3&gt;
&lt;p&gt;Conditional DiD relies on two core assumptions: conditional parallel trends and no anticipation, plus an overlap condition.&lt;/p&gt;







&lt;div class=&#34;math-environment assumption&#34; id=&#34;assumption-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Assumption 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(CausalML Book, page 457)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602152200078.png&#34; alt=&#34;image-20250602152200078&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;How to understand the &lt;strong&gt;overlap condition&lt;/strong&gt; (16.3.3)? It essentially imposes that there are control observations available for every value of $X$.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;the-key-insight-transformation-to-cross-sectional-problem&#34;&gt;The Key Insight: Transformation to Cross-Sectional Problem&lt;/h3&gt;
&lt;p&gt;By taking the difference,&lt;/p&gt;
&lt;p&gt;$$
\Delta Y = Y_{\text{after}} - Y_{\text{before}}
$$&lt;/p&gt;
&lt;p&gt;we transform panel data into a cross-sectional problem. This allows us to apply the same doubly robust framework used for cross-sectional ATT.&lt;/p&gt;
&lt;h3 id=&#34;the-unified-result&#34;&gt;The Unified Result&lt;/h3&gt;
&lt;p&gt;The Neyman orthogonal score for conditional DiD is &lt;strong&gt;identical&lt;/strong&gt; to the cross-sectional ATT score, where the outcome variable is simply the difference $\Delta Y$.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Neyman orthogonal score for ATT in conditional DiD&lt;/p&gt;







&lt;div class=&#34;math-environment proposition&#34; id=&#34;proposition-1&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Proposition 1&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(see CausalML Book)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602153412853.png&#34; alt=&#34;image-20250602153412853&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Neyman orthogonal score for ATT in cross-sectional setting&lt;/p&gt;







&lt;div class=&#34;math-environment proposition&#34; id=&#34;proposition-2&#34;&gt;
  
  &lt;div class=&#34;math-env-title&#34;&gt;
    &lt;strong&gt;Proposition 2&lt;/strong&gt; &lt;span class=&#34;math-env-name&#34;&gt;(see CausalML Book)&lt;/span&gt;.
  &lt;/div&gt;
  &lt;div class=&#34;math-env-content&#34;&gt;
    
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602153919836.png&#34; alt=&#34;image-20250602153919836&#34; style=&#34;zoom:40%;&#34; /&gt; &lt;br&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250602154015540.png&#34; alt=&#34;image-20250602154015540&#34; style=&#34;zoom:40%;&#34; /&gt;

  &lt;/div&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;mark&gt;Comparing to the score for the ATT in cross-sectional setting, we see that DiD score is &lt;strong&gt;identical&lt;/strong&gt; to that for learning the ATT under &lt;strong&gt;unconfoundedness&lt;/strong&gt; where the outcome variable is simply defined as $\Delta Y$ &lt;/mark&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    This elegant connection demonstrates that the doubly robust estimator for conditional DiD is equivalent to the doubly robust ATT estimator applied to the differenced outcome $\Delta Y$.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;p&gt;Chernozhukov, Victor, Christian Hansen, Nathan Kallus, Martin Spindler, and Vasilis Syrgkanis (2024), “Applied causal inference powered by ML and AI.”&lt;/p&gt;
&lt;p&gt;Ding, P. (2024). A First Course in Causal Inference. CRC Press.&lt;/p&gt;
&lt;p&gt;Callaway, Brantly and Pedro H. C. Sant’Anna (2021), “Difference-in-Differences with multiple time periods,” &lt;i&gt;Journal of Econometrics&lt;/i&gt;, Themed Issue: Treatment Effect 1, 225 (2), 200–230.&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K Newey, and Rahul Singh (2022), “Debiased machine learning of global and local parameters using regularized Riesz representers,” &lt;i&gt;The Econometrics Journal&lt;/i&gt;, 25 (3), 576–601.&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Whitney K. Newey, Victor Quintas-Martinez, and Vasilis Syrgkanis (2024), “Automatic debiased machine learning via riesz regression.”&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Big Picture of Debiased Machine Learning</title>
      <link>https://chenxing.space/blog/big-picture-of-debiased-machine-learning/</link>
      <pubDate>Tue, 25 Mar 2025 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/big-picture-of-debiased-machine-learning/</guid>
      <description>&lt;p&gt;Debiased machine learning (DML) is a generic recipe. The idea behind it is adding a correction term to the plug-in estimator of the functional, which leads to properties such as semi-parametric efﬁciency, double robustness, and Neyman orthogonality.&lt;/p&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250325151559540.png&#34; alt=&#34;image-20250325151559540&#34; style=&#34;zoom:80%;&#34; /&gt;
&lt;p&gt;(Auto)-DML is a &lt;strong&gt;Method-of-Moments estimator&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;debiased/orthogonal moment scores&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Why it matters?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;try to solve: model selection and/or &lt;strong&gt;regularization bias&lt;/strong&gt; from ML learners (e.g. Lasso)&lt;/li&gt;
&lt;li&gt;Neyman orthogonality: ensure the parameter of interest insentitive to first order perturbation of nuisance estimation&lt;/li&gt;
&lt;li&gt;double robustness&lt;/li&gt;
&lt;li&gt;asymptotic normality&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Key Idea: Debiasing is achieved by adding a correction term to the plug-in estimator of the functional&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Three representations: $\theta = \mathbb{E}[m(W,g)] = \mathbb{E}[Y\alpha(W)] = \mathbb{E}[g(W)\alpha(W)]$, where&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$g()$ is outcome regression;&lt;/li&gt;
&lt;li&gt;$\alpha()$ is Rieze Representer (RR);&lt;/li&gt;
&lt;li&gt;$m()$ is a continuous linear functional;&lt;/li&gt;
&lt;li&gt;$W = (D, X)$ is data containing treatment $D$ and covariates $X$;&lt;/li&gt;
&lt;li&gt;$Y(d)$ is  potential outcome&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Correct the residual using RR&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;mark&gt;$\mathbb{E}\{m(W,g) - \theta + \alpha(W)[Y-g(W)]\} = 0$&lt;/mark&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How to construct orthogonal moment function?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;orthogonal moment function = identifying moment function + first step influence function (FSIF)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;identifying moment function: $m(W,g) - \theta$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;involving &lt;strong&gt;outcome regression&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;FSIF: $\alpha(W)[Y-g(W)]$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;correct the residual using Rieze Representer (RR)&lt;/li&gt;
&lt;li&gt;Rieze Representer (RR)
&lt;ul&gt;
&lt;li&gt;In the case of ATE with binary treatment, RR are inverse propensity score terms&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;RR can be automatically characterized&lt;/strong&gt;; NO NEED to know its analytical form&lt;/li&gt;
&lt;li&gt;Can use random forests and NNet learners of RR&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Double Robustness&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;mark&gt;$\mathbb{E}[m(W ; g) -\theta_0  \left.+\alpha(W)(Y-g(W))\right] =-\mathbb{E}\left[\left(\alpha-\alpha_0\right)\left(g-g_0\right)\right]$&lt;/mark&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The score will be zero in expectation when &lt;strong&gt;either&lt;/strong&gt; $\alpha(W) = \alpha_0(W)$ &lt;strong&gt;or&lt;/strong&gt; $g(W) = g_0(W)$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Cross-fitting&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Why it matters?
&lt;ul&gt;
&lt;li&gt;Reduce overfitting bias&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>Study Notes on Bounding OVB 🌀 in Causal ML</title>
      <link>https://chenxing.space/blog/notes-bound-ovb-in-causal-ml/</link>
      <pubDate>Thu, 12 Sep 2024 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/notes-bound-ovb-in-causal-ml/</guid>
      <description>&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;p&gt;In empirical research, one of the challenges to causal inference is the potential presence of unobserved confounding. Even when we adjust for a wide range of observed covariates, there&amp;rsquo;s often a lingering concern: what if there are important variables we&amp;rsquo;ve failed to measure or include? This challenge necessitates a careful approach to sensitivity analysis, where we assess how strong unobserved confounders would need to be to meaningfully alter our conclusions.&lt;/p&gt;
&lt;figure&gt;
    &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/latest.png&#34; 
         alt=&#34;image of Mafūba technique&#34; 
         style=&#34;zoom:80%;&#34; 
         title=&#34;Mafūba sealing technique&#34; /&gt;
    &lt;figcaption&gt;
        &lt;strong&gt;Mafūba&lt;/strong&gt;: A technique designed to seal &lt;a href=&#34;https://dragonball.fandom.com/wiki/Demon&#34; target=&#34;_blank&#34;&gt;demons&lt;/a&gt; by drawing them into a container, which is then secured with a special &#34;Demon Seal&#34; &lt;a href=&#34;https://en.wikipedia.org/wiki/O-fuda&#34; target=&#34;_blank&#34;&gt;ofuda&lt;/a&gt;. &lt;cite&gt;Source: Image from &lt;a href=&#34;https://dragonball.fandom.com/wiki/Evil_Containment_Wave&#34; target=&#34;_blank&#34;&gt;Dragon Ball Fandom&lt;/a&gt;&lt;/cite&gt;
    &lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id=&#34;questions1&#34;&gt;Questions&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a href=&#34;#fn:1&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;How &amp;ldquo;&lt;strong&gt;strong&lt;/strong&gt;&amp;rdquo; would a particular confounder (or group of confounders) have to be to &lt;strong&gt;change the conclusions of a study&lt;/strong&gt;?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In a &lt;strong&gt;worst case scenario&lt;/strong&gt; how vulnerable is the study&amp;rsquo;s result to many or all unobserved confounders &lt;strong&gt;acting together&lt;/strong&gt;, possibly &lt;strong&gt;nonlinearly&lt;/strong&gt;?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Are these confounders or scenarios &lt;strong&gt;plausible&lt;/strong&gt;? How strong would they have to be &lt;strong&gt;relative to observed covariates (e.g female)&lt;/strong&gt;, in order to be problematic?&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How can we present these sensitivity results &lt;strong&gt;concisely&lt;/strong&gt; for &lt;strong&gt;easy routine reporting&lt;/strong&gt;?&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;tldr&#34;&gt;TL;DR&lt;/h2&gt;
&lt;p&gt;&lt;a href=&#34;https://doi.org/10.48550/arXiv.2112.13398&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Chernozhukov et al. (2024)&lt;/a&gt; introduces a framework for bounding omitted variable bias (OVB) in causal machine learning setting. They develop a general theory applicable to a broad class of causal parameters that can be expressed as linear functionals of the conditional expectation function. The key insight is to characterize the bias using the Riesz representation. Specifically, they &lt;strong&gt;express the OVB as the covariance of two error terms&lt;/strong&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;The error in the &lt;strong&gt;outcome regression&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The error in the &lt;strong&gt;Riesz representer (RR)&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This formulation leads to a bound on the squared bias that is the product of two terms:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The MSE of the outcome regression&lt;/li&gt;
&lt;li&gt;The MSE of the RR&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;theoretical-details&#34;&gt;Theoretical Details&lt;/h2&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    &lt;p&gt;Recall that:&lt;/p&gt;
&lt;p&gt;For classical linear regression, &amp;ldquo;Short equals long plus the effect of omitted times the regression of omitted on included.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&amp;ndash; Angrist and Pischke, Mostly Harmless Econometrics.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;in-linear-regression-setting&#34;&gt;In Linear Regression Setting&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Linear model:
$$
Y_i = \alpha + \beta D_i + \gamma^{\intercal} X_i + \delta U_i + \epsilon_i,
$$
where $U_i$ is an unobserved (scalar) confounding. Note that, if there are multiple confounders, we can regard $U_i$ as a function of them, which are summarized as a single confounder.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Recall the &lt;strong&gt;omitted variable bias formula&lt;/strong&gt;:
$$
\hat{\beta} \overset{p}{\rightarrow} \beta + \delta \times \frac{\mathrm{Cov}(U_i^{\perp X}, D_i^{\perp X})}{\mathbb{V}(D_i^{\perp X})},
$$&lt;/p&gt;
&lt;p&gt;where&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$D_i^{\perp X} = D_i - \mathrm{lm}(D\sim X)$ is the residual term&lt;/li&gt;
&lt;li&gt;$U_i^{\perp X}$ is the residual from $\mathrm{lm}(U\sim X)$&lt;/li&gt;
&lt;li&gt;$\frac{\mathrm{Cov}(U_i^{\perp X}, D_i^{\perp X})}{\mathbb{V}(D_i^{\perp X})}$ is the coefficient of residual-on-residual regression. It means how much variation of $U$ can be explained by $D$ condition on $X$. In other words, the association between $U$ and $D$ condition on $X$&lt;/li&gt;
&lt;li&gt;$\delta$ is the association between $Y$ and $U$ condition on $\{D, X\}$&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;How to bound the bias?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250129075223109.png&#34; alt=&#34;image-20250129075223109&#34; style=&#34;zoom:40%;&#34; /&gt;
&lt;/li&gt;
&lt;li&gt;Partial $R^2$ represents how much variation of $Y$ can explained by $U$ condition on $\{T, X\}$&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;in-causal-ml-setting&#34;&gt;In Causal ML Setting&lt;/h3&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;In the ideal case:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&amp;ldquo;long&amp;rdquo; regression: $Y = \theta D + f(X, A) + \epsilon$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Define $W := (D, X, A)$ as the &amp;ldquo;long&amp;rdquo; list of regressors&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$A$ is &lt;strong&gt;unobserved vector of covariates&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Assume $\mathbb{E}(\epsilon \mid D, X, A) = 0$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In the practical case:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&amp;ldquo;short&amp;rdquo; regression: $Y = \theta_sD + f_s(X) + \epsilon_s$&lt;/li&gt;
&lt;li&gt;Define $W_s := (D,X)$ as the “short” list of observed regressors because $A$ is unmeasured/unobservable&lt;/li&gt;
&lt;li&gt;Use $\theta_s$ to approximate $\theta$; need to bound $\theta_s - \theta$&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Conditional Expectation (outcome regression function)&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&amp;ldquo;long&amp;rdquo; CET,  $g(W) := \mathbb{E}(Y \mid D, X, A) = \theta D + f(X, A)$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&amp;ldquo;short&amp;rdquo; CET,  $g_s(W) := \theta_s D + f_s(X)$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;R&lt;/strong&gt;iesz &lt;strong&gt;R&lt;/strong&gt;epresenters (&lt;strong&gt;RR&lt;/strong&gt;):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&amp;ldquo;long&amp;rdquo; RR:
$$
\alpha:= \alpha(W):=\frac{D-\mathrm{E}[D \mid X, A]}{\mathrm{E}(D-\mathrm{E}[D \mid X, A])^2}
$$&lt;/p&gt;
&lt;p&gt;By FWL Theorem, we have
$$
\theta = \frac{Cov(Y\tilde{D})}{Var(\tilde{D})} = \mathbb{E}(y\alpha), \ \text{where } \tilde{D} = D-\mathrm{E}[D \mid X, A],
$$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&amp;ldquo;short&amp;rdquo; RR
$$
\alpha_s := \alpha_s\left(W^S\right):=\frac{D-\mathrm{E}[D \mid X]}{\mathrm{E}(D-\mathrm{E}[D \mid X])^2}
$$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;In special case, one can show:
$$
\alpha(W)=\frac{D}{P(D=1 \mid X, A)}-\frac{1-D}{P(D=0 \mid X, A)},
$$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;$$
\alpha_s(W)=\frac{D}{P(D=1 \mid X)}-\frac{1-D}{P(D=0 \mid X)},
$$&lt;/p&gt;
&lt;p&gt;&lt;mark&gt;&lt;strong&gt;The RR &amp;ldquo;looks like&amp;rdquo; the weights in Inverse Propensity Score weighting (IPW)&lt;/strong&gt;.&lt;/mark&gt; For example,&lt;/p&gt;

     $$
     \begin{aligned}
      \alpha_s(W) &amp;= \frac{D}{P(D=1 \mid X)}-\frac{1-D}{P(D=0 \mid X)} \\&amp; = \frac{D}{e(X)} - \frac{1-D}{1-e(X)}, \ \, \text{where } e(X) = P(D=1 \mid X)
      \end{aligned}
     $$
     
&lt;p&gt;Recall the property of IPW estimator, we have:&lt;/p&gt;

     $$
     \theta_s \overset{ipw}{=} \mathbb{E}\left[\frac{YD}{e(X)} - \frac{Y(1-D)}{1-e(X)}\right] = \mathbb{E}(Y\alpha_s)
     $$
     
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Note that, $\theta = \mathbb{E}(g \alpha)$. Why?

   $$
   \begin{aligned}
   \mathbb{E}(g \alpha) &amp; = \mathbb{E}\left\{ \mathbb{E}(Y \mid D, X, A) \frac{D-\mathrm{E}[D \mid X, A]}{\mathrm{E}(D-\mathrm{E}[D \mid X, A])^2} \right\} \\
   &amp;= \mathbb{E}\left\{ [\theta D + f(X, A)] \frac{D-\mathrm{E}[D \mid X, A]}{\mathrm{E}(D-\mathrm{E}[D \mid X, A])^2} \right\} \\
   &amp;= \theta 
   \end{aligned}
   $$
   &lt;/p&gt;
&lt;p&gt;For the third equation, we use the fact that: $\mathrm{E}[D \mid X, A] \ {\perp \!\!\! \perp} \ D-\mathrm{E}[D \mid X, A]$ and $f(X, A) \ {\perp \!\!\! \perp} \ D-\mathrm{E}[D \mid X, A]$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    &lt;p&gt;Short Summary:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$g$ is the outcome regression function&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\alpha$ is the &lt;strong&gt;R&lt;/strong&gt;iesz &lt;strong&gt;R&lt;/strong&gt;epresenter&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;We have $\theta = \mathbb{E}(Y\alpha) = \mathbb{E}(g \alpha)$ and $\theta_s = \mathbb{E}(Y\alpha_s) = \mathbb{E}(g_s \alpha_s)$.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;key-results&#34;&gt;Key Results&lt;/h3&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/image-20240913170504351.png&#34; alt=&#34;image-20240913170504351&#34; style=&#34;zoom:60%;&#34; /&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    &lt;p&gt;Key observation:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;OVB is bounded by $\mathbb{E}(\text{outcome regression error} \cdot \text{RR error})$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Square bias is bounded by $\text{MSE(regression)}  \cdot \text{MSE(RR)}$&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Idea of Proof:&lt;/strong&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Write $\theta = \mathbb{E}(g \alpha)$ and $\theta_s = \mathbb{E}(g_s \alpha_s)$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Use the fact that $\alpha_s \ {\perp \!\!\! \perp} \  g-g_s$ and $g_s \ {\perp \!\!\! \perp} \ \alpha - \alpha_s$&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Now, we know the omitted variable bias,&lt;/p&gt;
&lt;p&gt;$$\theta_s - \theta = \mathbb{E}(g_s - g)(\alpha_s - \alpha),$$&lt;/p&gt;
&lt;p&gt;is the &lt;strong&gt;covariance&lt;/strong&gt; between the regression error,&lt;/p&gt;
&lt;p&gt;$$g_s - g = \mathbb{E}(Y|W_s) - \mathbb{E}(Y|W),$$&lt;/p&gt;
&lt;p&gt;and RR error,&lt;/p&gt;
&lt;p&gt;$$\alpha_s - \alpha,$$&lt;/p&gt;
&lt;p&gt;that is, the &amp;ldquo;weights in IPW&amp;rdquo; calculated using the &lt;em&gt;short&lt;/em&gt; list of regressors, minus the &amp;ldquo;weights in IPW&amp;rdquo; using the &lt;em&gt;long&lt;/em&gt; list of regressors.&lt;/p&gt;
&lt;p&gt;As we can see, this is not easy to interpret, so the following Corollary provides a more intuitive understanding using the familiar $R^2$.&lt;/p&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/image-20240913170858030.png&#34; alt=&#34;image-20240913170858030&#34; style=&#34;zoom:60%;&#34; /&gt;
&lt;p&gt;Note that,&lt;/p&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/image-20250118105142237.png&#34; alt=&#34;image-20250118105142237&#34; style=&#34;zoom:10%;&#34; /&gt;
&lt;p&gt;The bias is thus bounded by the additional variation that omitted variables cause in the regression function and Riezs representers.&lt;/p&gt;
&lt;p&gt;We can make this easier to interpret with further algebra,&lt;/p&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/image-20250118105610894.png&#34; alt=&#34;image-20250118105610894&#34; style=&#34;zoom:10%;&#34; /&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    &lt;p&gt;Short Summary:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;$S$ is the scale of the bias, which can be identified from the observed data.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The confounding strength $C_Y$ and $C_D$ have to be restricted by the analyst.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$C_{Y}^2$ measures the &lt;strong&gt;proportion of residual variation of the outcome&lt;/strong&gt; explained by latent confounders; in short, &lt;mark&gt;the strength of unmeasured confounding in outcome equation&lt;/mark&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$C_{D}^2$ measures the &lt;strong&gt;proportion of residual variation of the treatment&lt;/strong&gt; explained by latent confounders; in short, &lt;mark&gt;the strength of unmeasured confounding in treatment equation&lt;/mark&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;figure&gt;
    &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/image-20240913193118914.png&#34; 
         alt=&#34;image-20240913193118914&#34; 
         style=&#34;zoom:50%;&#34; 
         title=&#34;Big Picture&#34; /&gt;
    &lt;figcaption&gt;
        &lt;cite&gt;Source: &lt;a href=&#34;https://www.youtube.com/watch?v=PQtYqKfxH_I&#34; target=&#34;_blank&#34;&gt;Victor’s tutorial at the Chamberlain Seminar&lt;/a&gt;&lt;/cite&gt;
    &lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id=&#34;empirical-challenge&#34;&gt;Empirical Challenge&lt;/h2&gt;
&lt;p&gt;How do we determine the plausible strength of these unobserved confounders?&lt;/p&gt;
&lt;div class=&#34;alert alert-note&#34;&gt;
  &lt;div&gt;
    How to set the strength of these unobserved confounders? $C_{Y}^2 = ?$ and $C_{D}^2 = ?$
  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;This is a critical question, as setting these values arbitrarily could lead to overly conservative or overly optimistic interpretations of our results. We need a principled approach to guide our choices.&lt;/p&gt;
&lt;h3 id=&#34;benchmarking&#34;&gt;Benchmarking&lt;/h3&gt;
&lt;p&gt;One solution is &lt;strong&gt;benchmarking&lt;/strong&gt;. This approach leverages the observed data to inform our judgments about unobserved confounders. The basic idea is the following: &lt;strong&gt;we can use the impact of observed covariates as a reference point for the potential impact of unobserved ones&lt;/strong&gt;. For instance, if we&amp;rsquo;ve measured income, the &amp;ldquo;most&amp;rdquo; important factor, and found it explains 15% of the variation in our outcome, we might reason that an unobserved confounder is unlikely to have an even larger effect.&lt;/p&gt;
&lt;p&gt;In practice, benchmarking can take several forms. We might purposely omit a known important covariate, refit our model, and observe the change. This gives us a concrete example of how omitting an important variable affects our estimates. Alternatively, we could express the strength of unobserved confounders relative to observed ones. For example, we might consider scenarios where an unobserved confounder is as strong as income, or perhaps 25% as strong.&lt;/p&gt;
&lt;p&gt;Recent research (Chernozhukov et al.,2024) has demonstrated the utility of this approach. In a study on 401(k) eligibility, they used the observed impact of income, IRA participation, and two-earner status as benchmarks for potential unobserved firm characteristics. Similarly, in a study on gasoline demand, the known impact of income brackets informed the choice of sensitivity parameters for potential remnant income effects.&lt;/p&gt;
&lt;h3 id=&#34;robustness-value&#34;&gt;Robustness Value&lt;/h3&gt;
&lt;p&gt;Another useful concept is the &lt;strong&gt;Robustness Value (RV)&lt;/strong&gt; which represents the minimum strength of confounding required to change a study&amp;rsquo;s conclusions. This provides a clear threshold for evaluating the robustness of results.&lt;/p&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/image-20250119083624670.png&#34; alt=&#34;image-20250119083624670&#34; style=&#34;zoom:67%;&#34; /&gt;
&lt;p&gt;The idea of robustness values is to quickly communicate how robust the &lt;em&gt;short estimate&lt;/em&gt; is to &lt;strong&gt;systematic errors&lt;/strong&gt; due to residual confounding. For example, $RV_{\theta = 0, \alpha = 0.05}$ measures the minimal strength on both confounding factors such that the estimated confidence bound for the ATE would include zero, at the 5% significance level.&lt;/p&gt;
&lt;h4 id=&#34;401k-example-in-the-paper&#34;&gt;401(k) example in the paper&lt;/h4&gt;
&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/image-20250119090152561.png&#34; alt=&#34;image-20250119090152561&#34; style=&#34;zoom:50%;&#34; /&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Table 1 illustrates our proposal for a minimal sensitivity reporting of causal effect estimates. Beyond the usual estimates under the assumption of conditional ignorability, it reports the robustness values of the short estimate. &lt;br&gt;
Starting with the PLM, the $RV_{\theta = 0, \alpha = 0.05} = 5.4 \%$ means that unobserved confounders that explain less than 5.4% of the residual variation, &lt;strong&gt;both&lt;/strong&gt; of the treatment, and of the outcome, are not sufficiently strong to bring the lower limit of the confidence bound to zero, at the 5% significance level.&amp;rdquo; &lt;/br&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Interpretation of RV:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;A &lt;strong&gt;higher RV&lt;/strong&gt; suggests that the causal estimate is &lt;strong&gt;more robust&lt;/strong&gt; to omitted variable bias because stronger confounding would be needed to alter the conclusions.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;A lower RV suggests that a relatively small amount of confounding could change the conclusions.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In summary, the robustness value (RV) provides a concise way to communicate how sensitive the results of a causal analysis are to potential omitted variable bias. It is the minimum strength of confounding that would be required to change the conclusions. A &lt;strong&gt;larger RV indicates a more robust analysis&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&#34;software-to-implement&#34;&gt;Software to Implement&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/carloscinelli/dml.sensemakr&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;carloscinelli/dml.sensemakr: Sensitivity analysis tools for causal ML&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=p3dfHj6ki68&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;useR! 2020: sensemakr: Sensitivity Analysis Tools for OLS (C.Cinelli), regular - YouTube&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;p&gt;Chernozhukov, Victor, Carlos Cinelli, Whitney Newey, Amit Sharma, and Vasilis Syrgkanis. “Long Story Short: Omitted Variable Bias in Causal Machine Learning.” arXiv, May 26, 2024. &lt;a href=&#34;https://doi.org/10.48550/arXiv.2112.13398&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.48550/arXiv.2112.13398&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://docs.doubleml.org/stable/guide/sensitivity.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;10. Sensitivity analysis&lt;/a&gt; in DoubleML package&lt;/p&gt;
&lt;p&gt;Angrist, Joshua D., and Jörn-Steffen Pischke. Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton University Press, 2009.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://carloscinelli.com/files/Cinelli%20et%20al%20%282020%29%20-%20sensemakr.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;sensemakr: Sensitivity Analysis Tools for OLS in R and Stata&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Carlos Cinelli and Chad Hazlett. ‘Making sense of sensitivity: Extending omitted variable bias’. In: Journal of the Royal Statistical Society: Series B (Statistical Methodology) 82.1 (2020), pp. 39–67 (cited on pages 23, 26).&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=FAJe0rNNjBs&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;sensitivity analysis lecture&lt;/a&gt;&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;These questions are in the slides of &lt;a href=&#34;https://www.youtube.com/watch?v=p3dfHj6ki68&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;this tutorial&lt;/a&gt;&amp;#160;&lt;a href=&#34;#fnref:1&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Notes on Causal Survival Forest 🌲⏳</title>
      <link>https://chenxing.space/blog/causal-survival-forest-notes/</link>
      <pubDate>Thu, 05 Sep 2024 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/causal-survival-forest-notes/</guid>
      <description>&lt;p&gt;In this post, I provide summary notes on the paper &amp;ldquo;Estimating Heterogeneous Treatment Effects with Right-Censored Data via Causal Survival Forests&amp;rdquo; by Cui et al. (2023).&lt;/p&gt;
&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;p&gt;&lt;mark&gt;How to estimate heterogeneous treatment effects with right-censored data?&lt;/mark&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneous treatment effect (HTE) estimation&lt;/strong&gt; plays a central role in data-driven personalization&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Existing methods often can&amp;rsquo;t handle &lt;strong&gt;&lt;mark&gt;censored survival outcomes&lt;/mark&gt;&lt;/strong&gt;, common in medical/business applications&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;causal-survival-forests-csf&#34;&gt;Causal Survival Forests (CSF)&lt;/h2&gt;
&lt;p&gt;To address this challenge, the paper proposes causal survival forests (CSF)&lt;/p&gt;
&lt;center&gt;&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/1233.jpg&#34; alt=&#34;1233&#34; style=&#34;zoom:50%;&#34; /&gt;&lt;/center&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;An adaptation of the causal forest algorithm of Athey et al. (2019)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;It adjusts for censoring using doubly robust estimating equations developed in the survival analysis literature&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Advantages&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;mark&gt;&lt;strong&gt;Doubly robust&lt;/strong&gt;&lt;/mark&gt;, computationally tractable, and outperforms available baselines in our experiments&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Good statistical properties &amp;ndash; UCAN&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;setting-and-notation&#34;&gt;Setting and notation&lt;/h2&gt;
&lt;p&gt;Assume i.i.d tuples $\{X_i, T_i, C_i, W_i\}$, where&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;$X_i \in \mathcal{X}$ denote &lt;strong&gt;covariates&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;$T_i \in \mathbb{R}_{+}$ is the &lt;strong&gt;survival time&lt;/strong&gt; for $i$th unit&lt;/li&gt;
&lt;li&gt;$C_i \in \mathbb{R}_{+}$ is the &lt;strong&gt;censoring time&lt;/strong&gt; (the time at which $i$th unit gets censored)&lt;/li&gt;
&lt;li&gt;$W_i \in \{0,1\}$  denotes a &lt;strong&gt;binary treatment&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Using potential outcome framework, posit potential outcomes $\{T_i(1), T_i(0)\}$ s.t. $T_i = T_i(W_i)$, we need to estimate the conditional average treatment effect (CATE)&lt;/p&gt;
&lt;p&gt;$$
\tau(x)=\mathbb{E}\left[y(T_i(1)) - y(T_i(0)) \mid X_i=x\right],
$$&lt;/p&gt;
&lt;p&gt;where $y()$ is the outcome transformation. For example,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$y(T) = T$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$y(T) = T \wedge h$ for the &lt;strong&gt;restricted mean survival time (RMST)&lt;/strong&gt;; here $h$ is some chosen maximum considered time&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$y(T) = \1\{T \ge h\}$ for the &lt;strong&gt;survival probability&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;mark&gt;To estimate $\tau(x)$, the main challenge is that $T_i$ is not always observable. We &lt;strong&gt;can only observe&lt;/strong&gt;&lt;/mark&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;censored survival time&lt;/strong&gt;: $U_i = T_i \wedge C_i$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;non-censoring indicator&lt;/strong&gt;: $\Delta_i = 1\{T_i \le C_i\}$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Based on Assumption 1 (see later), we define the &lt;strong&gt;effective non-censoring indicator&lt;/strong&gt; as follows:&lt;/p&gt;

\begin{align}
\Delta_i^h &amp;= 1\{T_i \wedge h \le C_i\} \\
           &amp;\overset{(2)}{=} \Delta_i \vee 1\{U_i \ge h\}
\end{align}

&lt;p&gt;Note that, for the eqn (2), everything is observed. We can regard an observation with $\Delta_i^h = 1$ as a &lt;strong&gt;complete observation&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&#34;assumptions&#34;&gt;Assumptions&lt;/h2&gt;
&lt;p&gt;In order to identify treatment effects, we need to rely on two sets of assumptions.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Assumption 2-4 enable us to identify the causal effect of $W_i$ on $T_i$ without censoring&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Assumption 5-6 is to guarantee that censoring due to $C_i$ does not break identification results&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Assumption 1 (Finite Horizon)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;$$
y(t) = y(h), \quad \forall \ t \ge h, \ 0&amp;lt;h&amp;lt;\infty
$$&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Assumption 2 (Potential Outcomes)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;$$
\{T_i(1), T_i(0)\} \quad s.t. \quad T_i = T_i(W_i) \quad a.s.
$$ 
&lt;strong&gt;Assumption 3 (Ignorability)&lt;/strong&gt;&lt;/p&gt;
$$
\{T_i(1), T_i(0)\} \perp W_i \mid X_i
$$ 
&lt;p&gt;&lt;strong&gt;Assumption 4 (Overlap)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Propensity score $e(x) = \mathbb{P}(W_i = 1\mid X_i = x)$ is uniformly bounded away from 0 and 1,&lt;/p&gt;
&lt;p&gt;$$
\eta_e \le e(x)\le 1- \eta_e, \quad  0&amp;lt; \eta_e \le \frac{1}{2}
$$&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Assumption 5 (Ignorable censoring)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Censoring is independent of survival time conditionally on treatment and covariates,&lt;/p&gt;
&lt;p&gt;$$
T_i \perp C_i \mid W_i, X_i
$$&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Assumption 6 (Positivity)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;$$
\mathbb{P}(C_i &amp;lt; h | W_i, X_i) \le 1- \eta_c, \quad 0&amp;lt;\eta_c\le1
$$&lt;/p&gt;
&lt;h2 id=&#34;causal-forests-without-censoring&#34;&gt;Causal Forests Without Censoring&lt;/h2&gt;
&lt;p&gt;How does causal forest work?&lt;/p&gt;
&lt;p&gt;Essentially, we are running a “forest”-localized version of Robinson’s regression&lt;/p&gt;
&lt;p&gt;$$
\tau(x):=\operatorname{lm}\left(Y_i-\hat{m}^{(-i)}\left(X_i\right) \sim W_i-\hat{e}^{(-i)}\left(X_i\right), \text { weights }=\textcolor{blue}{\alpha_i(x)}\right),
$$&lt;/p&gt;
&lt;p&gt;where&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$\textcolor{blue}{\alpha_i(x)}$ capture how &amp;ldquo;similar&amp;rdquo; a target sample $x$ is to each of the training samples $X_i$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\hat{m}$ and $\hat{e}$ are machine learning estimates for the outcome and propensity score models with cross-fitting&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Using notations in previous section, we estimate $\tau(x)$ by solving the following equation,&lt;/p&gt;
&lt;p&gt;$$
\sum_{i=1}^n \alpha_i(x) \psi_{\hat{\tau}(x)}^{(c)}\left(X_i, y\left(T_i\right), W_i ; \hat{e}, \hat{m}\right)=0 \tag{cf}
$$&lt;/p&gt;
&lt;p&gt;where,&lt;/p&gt;

\begin{aligned}
\psi_\tau^{(c)}\left(X_i, y\left(T_i\right), W_i ; \hat{e}, \hat{m}\right) &amp;= \left[ W_i - \hat{e}\left(X_i\right) \right] \quad \times\\
&amp; \left[ y\left(T_i\right) - \hat{m}\left(X_i\right) - \tau \left( W_i - \hat{e}\left(X_i\right) \right) \right]
\end{aligned}

&lt;p&gt;is the orthogonal &lt;strong&gt;complete&lt;/strong&gt; score function (shown as up-script $^{(c)}$),&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$e(x) = \mathbb{P}(W_i = 1\mid X_i = x)$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$m(x) = \mathbb{E}(y(T_i) \mid X_i)$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\hat{e}(X_i)$ and $\hat{m}(X_i)$ are estimates derived via cross-fitting&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;adjusting-for-censoring-via-weighting&#34;&gt;Adjusting for Censoring via Weighting&lt;/h2&gt;
&lt;p&gt;In the presence of censoring, the $T_i$ in equation (cf) is no longer observable.&lt;/p&gt;
&lt;p&gt;Simply ignoring censoring and building models on with complete observations (i.e. $\Delta_i^h = 1$) would lead to bias.&lt;/p&gt;
&lt;h2 id=&#34;simple-censoring-adjustment-via-ipcw&#34;&gt;Simple Censoring Adjustment via IPCW&lt;/h2&gt;
&lt;p&gt;Define the &lt;strong&gt;conditional survival function for censoring process&lt;/strong&gt; as
$$
S_w^C(s \mid x)=\mathbb{P}\left[C_i \geq s \mid W_i=w, X_i=x\right]
$$
We have,
$$
\mathbb{P}\left[\Delta_i^h=1 \mid X_i, W_i, T_i\right]=S_{W_i}^C\left(T_i \wedge h \mid X_i\right)
$$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;the LHS is the conditional probability of observing a complete observations (i.e. $\Delta_i^h = 1$)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;the RHS is the conditional probability that censoring time is greater than survival time&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Does the above $\mathbb{P}(\Delta_i^h = 1 \mid \cdots)$ look like propensity score function?&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The main idea of IPCW estimation is to &lt;strong&gt;only consider complete cases&lt;/strong&gt;, but &lt;strong&gt;up-weight all complete observations&lt;/strong&gt; by $1/S_{W_i}^C\left(T_i \wedge h \mid X_i\right)$ to compensate for censoring.&lt;/p&gt;
&lt;p&gt;As a result, IPCW estimators succeed in eliminating censoring bias.&lt;/p&gt;
&lt;p&gt;With IPCW, we estimate $\tau(x)$ by solving the following equation,&lt;/p&gt;
&lt;p&gt;$$
\sum_{\left\{i: \Delta_i^h=1\right\}} \frac{\alpha_i(x)}{\hat{S}_{W_i}^C\left(T_i \wedge h \mid X_i\right)} \psi_{\hat{\tau}(x)}^{(c)}\left(X_i, y\left(T_i\right), W_i ; \hat{e}, \hat{m}\right)=0 \tag{IPCW}
$$ 
Let&amp;rsquo;s compare the equation (cf) v.s (IPCW),&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;For eqn (cf), we sum over all observations; for eqn (IPCW), we only sum over complete observations&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;For eqn (IPCW), we add $\frac{1}{1/S_{W_i}^C\left(T_i \wedge h \mid X_i\right)}$ as a part of weight&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For more details on IPCW, please check:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Chapter 8 and 12 in the textbook &lt;i&gt;Causal Inference: What If&amp;quot;&lt;/i&gt; (Hernán and Robins, 2020). In particular, &amp;ldquo;Ch 12.6 Censoring and missing data&amp;rdquo; is very helpful.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Chapter 21 &amp;ldquo;Treatment Heterogeneity with Survival Outcomes&amp;rdquo; in the textbook &lt;i&gt;Handbook of Matching and Weighting Adjustments for Causal Inference&lt;/i&gt; (Zubizarreta et al., 2023)&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;a-doubly-robust-correction&#34;&gt;A Doubly Robust Correction&lt;/h2&gt;
&lt;p&gt;Two limitations of IPCW approach:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Only use complete observations; throw away all observations with $\Delta_i^h = 0$, and this may hurt efficiency&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;IPCW-type methods are generally not robust to estimation errors; Neyman orthogonality condition (Chernozhukov et al. 2018) does not hold&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id=&#34;csf-method&#34;&gt;CSF Method&lt;/h3&gt;
&lt;p&gt;CSF method does not rely on IPCW. Instead, it relies on a more robust approach to making estimating equations robust to censoring.&lt;/p&gt;
&lt;p&gt;Recall the simplest case (without censoring), we have,&lt;/p&gt;
&lt;p&gt;$$
\sum_{i=1}^n \alpha_i(x) \psi_{\hat{\tau}(x)}^{(c)}\left(X_i, y\left(T_i\right), W_i ; \hat{e}, \hat{m}\right)=0 \tag{cf}
$$&lt;/p&gt;
&lt;p&gt;where&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt; $\psi_{\hat{\tau}(x)}^{(c)} (\cdot)$  is the score function with &lt;mark&gt;&lt;strong&gt;the complete data&lt;/strong&gt;&lt;/mark&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;mark&gt;The idea of CSF is to convert the above (cf) equation into a &lt;strong&gt;censoring robust estimating equation&lt;/strong&gt; by using estimates of the survival and censoring processes.&lt;/p&gt;
&lt;p&gt;Now, we estimate the $\tau(x)$ by solving the following equation,&lt;/p&gt;
$$
\sum_{i=1}^n \alpha_i(x) \psi_{\hat{\tau}(x)}\left(X_i, y\left(U_i\right), U_i \wedge h, W_i, \Delta_i^h ; \hat{e}, \hat{m}, \hat{\lambda}_w^C, \hat{S}_w^C, \hat{Q}_w\right)=0,
$$ 
&lt;p&gt;where the score function is,&lt;/p&gt;
&lt;center&gt;&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/image-20240919213740735.png&#34; alt=&#34;image-20240919213740735&#34; style=&#34;zoom:50%;&#34; /&gt;&lt;/center&gt;
&lt;p&gt;the &lt;strong&gt;conditional expectation of the transformed survival time&lt;/strong&gt; is defined as:&lt;/p&gt;
&lt;p&gt;$$
Q_w(s \mid x)=\mathbb{E}\left[y\left(T_i\right) \mid X_i=x, W_i=w, T_i \wedge h&amp;gt;s\right]
$$&lt;/p&gt;
&lt;p&gt;and the associated &lt;strong&gt;conditional hazard function&lt;/strong&gt; is defined as:&lt;/p&gt;
&lt;p&gt;$$
\lambda_w^{\mathrm{C}}(s \mid x)=-\frac{d}{d s} \log S_w^{\mathrm{C}}(s \mid x)
$$&lt;/p&gt;
&lt;p&gt;$\hat{Q}_w(s \mid x), \hat{S}_w^C(s \mid x)$ and $\hat{\lambda}_w^C(s \mid x)$ are cross-fit nuisance parameter estimates.&lt;/p&gt;
&lt;mark&gt;How to understand the above score function?&lt;/mark&gt;
&lt;blockquote&gt;
&lt;p&gt;The short answer is that that functional form emerges for the math (i.e., the desire for a doubly robust adjustment); and, unlike with the basic AIPW formula, it&amp;rsquo;s not as immediately intuitive.&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a href=&#34;#fn:1&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2 id=&#34;-key-points&#34;&gt;💡 KEY Points&lt;/h2&gt;
&lt;p&gt;We should think about the &lt;strong&gt;Neyman-orthogonal property&lt;/strong&gt;. In summary, CSF alleviates the drawbacks of IPCW so by taking the (complete-data) causal forest estimating equation $\psi_{\tau(x)}^{(c)}(T, W, \ldots)$ (the &amp;ldquo;R-learner&amp;rdquo;) and turn it into a censoring robust estimating equation $\psi_{\tau(x)}(Y, W, \ldots)$ by using estimates of the &lt;strong&gt;survival&lt;/strong&gt; and &lt;strong&gt;censoring processes&lt;/strong&gt;&lt;sup id=&#34;fnref:2&#34;&gt;&lt;a href=&#34;#fn:2&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;2&lt;/a&gt;&lt;/sup&gt;. (&amp;quot;&amp;hellip;&amp;quot; refers to additional nuisance parameters):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;censoring process: $P\left[C_i&amp;gt;t \mid X_i=x, W_i=w\right]$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;survival process: $P\left[T_i&amp;gt;t \mid X_i=x, W_i=w\right]$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;mark&gt;The upshot of this &amp;ldquo;orthogonal&amp;rdquo; estimating equation is that it will be &lt;strong&gt;consistent if either the survival or censoring process is correctly specified&lt;/strong&gt;,&lt;/mark&gt; which is very beneficial when we want to estimate these by modern ML tools, such as random survival forests.&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    CSF approach is doubly robust in the sense that we can obtain the consistent estimator either the survival or censoring process is correctly specified.
  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;For more details, &lt;a href=&#34;https://www.degruyter.com/document/doi/10.2202/1557-4679.1052/html?lang=en&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Rubin &amp;amp; van der Laan (2007)&lt;/a&gt; and the chapter on RCTs with time-to-event data in &lt;a href=&#34;https://link.springer.com/chapter/10.1007/978-1-4419-9782-1_17&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Targeted Learning (2011)&lt;/a&gt; gives some more digestible details on doubly robust estimation with survival data.&lt;sup id=&#34;fnref:3&#34;&gt;&lt;a href=&#34;#fn:3&#34; class=&#34;footnote-ref&#34; role=&#34;doc-noteref&#34;&gt;3&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;I also found the &lt;a href=&#34;https://gist.github.com/erikcs/cb8325fefe8bdfad6fc230015ddfe9cb&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;code&lt;/a&gt; in &lt;code&gt;grf&lt;/code&gt; GitHub repo helpful to understand the implement procedure. Specifically, check the following lines:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# The conditional survival function for the survival process.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;sf.survival&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;do.call&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;grf&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;::&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;survival_forest&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;c&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;list&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;X&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;cbind&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;X&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;W&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;),&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;Y&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;Y&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;D&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;D&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;),&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;args.nuisance&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# The conditional survival function for the censoring process.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;sf.censor&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;do.call&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;grf&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;::&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;survival_forest&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;c&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;list&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;X&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;cbind&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;X&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;W&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;),&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;Y&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;Y&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;D&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;D&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;),&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;args.nuisance&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;p&gt;Athey, Susan, Julie Tibshirani, and Stefan Wager. 2019. “Generalized Random Forests.” &lt;i&gt;The Annals of Statistics&lt;/i&gt; 47 (2): 1148–78. &lt;a href=&#34;https://doi.org/10.1214/18-AOS1709&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1214/18-AOS1709&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Chernozhukov, Victor, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. 2018. “Double/Debiased Machine Learning for Treatment and Structural Parameters.” &lt;i&gt;The Econometrics Journal&lt;/i&gt; 21 (1): C1–68. &lt;a href=&#34;https://doi.org/10.1111/ectj.12097&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1111/ectj.12097&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Cui, Yifan, Michael R Kosorok, Erik Sverdrup, Stefan Wager, and Ruoqing Zhu. 2023. “Estimating Heterogeneous Treatment Effects with Right-Censored Data via Causal Survival Forests.” &lt;i&gt;Journal of the Royal Statistical Society Series B: Statistical Methodology&lt;/i&gt; 85 (2): 179–211. &lt;a href=&#34;https://doi.org/10.1093/jrsssb/qkac001&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1093/jrsssb/qkac001&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Hernán MA, Robins JM (2020). Causal Inference: What If. Boca Raton: Chapman &amp;amp; Hall/CRC.&lt;/p&gt;
&lt;p&gt;Zubizarreta, J. R., Stuart, E. A., Small, D. S., &amp;amp; Rosenbaum, P. R. (2023). &lt;i&gt;Handbook of Matching and Weighting Adjustments for Causal Inference&lt;/i&gt;. CRC Press.&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;This was suggested by Professor Wager in an email conversation.&amp;#160;&lt;a href=&#34;#fnref:1&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:2&#34;&gt;
&lt;p&gt;Check more on &lt;a href=&#34;https://grf-labs.github.io/grf/articles/survival.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;grf tutorial: Causal forest with time-to-event data&lt;/a&gt;&amp;#160;&lt;a href=&#34;#fnref:2&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:3&#34;&gt;
&lt;p&gt;Suggested by &lt;a href=&#34;https://sites.google.com/view/erikcs#h.uqgrvlridx4y&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Erik Sverdrup&lt;/a&gt;. Many thanks!&amp;#160;&lt;a href=&#34;#fnref:3&#34; class=&#34;footnote-backref&#34; role=&#34;doc-backlink&#34;&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>A walkthrough of how Causal Forest 🌲 works</title>
      <link>https://chenxing.space/blog/a-walkthrough-of-how-causal-forest-works/</link>
      <pubDate>Tue, 20 Aug 2024 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/a-walkthrough-of-how-causal-forest-works/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;In this post, I will go over how causal forest works based on the &lt;a href=&#34;https://grf-labs.github.io/grf/articles/grf_guide.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;tutorial in grf R package&lt;/a&gt;. Causal Forests offer a flexible, data-driven approach to estimating varied treatment effects, bridging machine learning and causal inference techniques.&lt;/p&gt;
&lt;h2 id=&#34;common-setting&#34;&gt;Common Setting&lt;/h2&gt;
&lt;p&gt;If we are working on an &lt;strong&gt;observational&lt;/strong&gt; study, we have the following data:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Outcome variable: $Y_i$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Binary treatment indicator: $W_i = \{0, 1\}$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;A set of covariates: $X_i$&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Let&amp;rsquo;s assume that the following conditions hold:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;mark&gt;(Assumption 1) $W_i$ is unconfounded given $X_i$ (i.e. treatment is as good as random given covariates).&lt;/mark&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$$\{Y_i(0), Y_i(1)\} \perp W_i | X_i$$&lt;/p&gt;
&lt;ol start=&#34;2&#34;&gt;
&lt;li&gt;
&lt;mark&gt;(Assumption 2) The confounders $X_i$ have a linear effect on $Y_i$.&lt;/mark&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;mark&gt;(Assumption 3) The treatment effect $\tau$ is constant.&lt;/mark&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Then we could run a regression of the type&lt;/p&gt;
&lt;p&gt;$$
Y_i = \tau W_i + \beta X_i + \epsilon_i
$$&lt;/p&gt;
&lt;p&gt;and interpret the estimate of $\hat{\tau}$ as the average treatment effect (ATE) $\tau = \mathbb{E}(Y_i(1) - Y_i(0))$.&lt;/p&gt;
&lt;h2 id=&#34;relaxing-assumptions&#34;&gt;Relaxing Assumptions&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Assumption 1&lt;/strong&gt; is an &amp;ldquo;identifying&amp;rdquo; assumption we have to live with&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Assumption 2&lt;/strong&gt; and &lt;strong&gt;Assumption 3&lt;/strong&gt; are modeling assumptions that we can question.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;relaxing-assumption-2-partially-linear-model-plr&#34;&gt;Relaxing Assumption 2: Partially Linear Model (PLR)&lt;/h3&gt;
&lt;p&gt;Assumption 2 is a strong parametric modeling assumption that requires the confounders to have a linear effect on the outcome, and that we should be able to relax by relying on semi-parametric statistics.&lt;/p&gt;
&lt;p&gt;We can instead posit the partially linear model:&lt;/p&gt;
&lt;p&gt;$$
Y_i = \tau W_i + f(X_i) + \epsilon_i, \ \ \ \mathbb{E}(\epsilon_i | X_i, W_i) = 0
$$&lt;/p&gt;
&lt;p&gt;How do we get around estimating $\tau$ when we do not know $f(X_i)$?&lt;/p&gt;
&lt;p&gt;Define the propensity score as
$$e(x) = \mathbb{E}(W_i | X_i = x),$$&lt;/p&gt;
&lt;p&gt;and the conditional mean of $Y$ as&lt;/p&gt;
&lt;p&gt;$$m(x) = \mathbb{E}(Y_i | X_i = x) = f(x) + \tau e(x).$$&lt;/p&gt;
&lt;p&gt;By Robinson (1988), we can rewrite the above equation in &lt;strong&gt;&amp;ldquo;centered&amp;rdquo;&lt;/strong&gt; form:&lt;/p&gt;
&lt;p&gt;$$Y_i - m(x) = \tau \cdot [W_i - e(x)]  + \epsilon_i$$&lt;/p&gt;
&lt;p&gt;This formulation has great practical appeal, as it means $\tau$ can be estimated by &lt;mark&gt;&lt;strong&gt;residual-on-residual regression&lt;/strong&gt;&lt;/mark&gt;.&lt;/p&gt;
&lt;p&gt;Good properties 😀:  Robinson (1988) shows that this approach yields root-n consistent estimates of $\tau$, even if estimates of $m(x)$ and $e(x)$ converge at a slower rate (&amp;ldquo;4-th root&amp;rdquo; in particular). This property is often referred to as &lt;strong&gt;orthogonality&lt;/strong&gt; and &lt;mark&gt;is a desirable property that essentially tells you that given noisy &amp;ldquo;nuisance&amp;rdquo; estimates ($m(x)$ and $e(x)$) you can still recover &amp;ldquo;good estimates of your target parameter ($\tau$).&lt;/mark&gt; For more details, please refer to Wager, Stefan “STATS 361: Causal Inference” Lecture 3.&lt;/p&gt;
&lt;p&gt;But how to estimate $m(x)$ and $e(x)$?&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Use modern machine learning models!&lt;/strong&gt; One could use boosting, random forest, and etc to estimate $m(x)$ and $e(x)$ because what we need is just &amp;ldquo;reasonable accurate&amp;rdquo; predictions, i.e.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;$$\mathbb{E}\left[(\hat{m}(X)-m(X))^2\right]^{\frac{1}{2}}, \mathbb{E}\left[(\hat{e}(X)-e(X))^2\right]^{\frac{1}{2}}=o_P\left(\frac{1}{n^{1 / 4}}\right)$$&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Issue with Direct Plug-in of Estimates&lt;/strong&gt;: Directly plugging in $\hat{m}(x)$ and $\hat{e}(x)$  into the residual-on-residual regression typically leads to bias because modern ML methods regularize to trade off bias and variance.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Solution via Cross-Fitting&lt;/strong&gt;: cross-fitting, where the prediction for observation $i$  is obtained without using unit  $i$  for estimation, can help overcome this bias (Chernozhukov et al. 2018).&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    Recap: We have a way to adopt the modern ML toolkit to &lt;em&gt;non-parametrically control for confounding&lt;/em&gt; when estimating an ATE, and still retain desirable statistical properties such as unbiased-ness and consistency.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;relaxing-assumption-3-non-constant-treatment-effects&#34;&gt;Relaxing Assumption 3: Non-constant treatment effects&lt;/h3&gt;
&lt;p&gt;Non-constant treatment effects occur when the impact of a treatment varies across different subgroups or individuals. This concept relaxes the assumption of homogeneous treatment effects, where the treatment is assumed to have the same impact on all units.&lt;/p&gt;
&lt;p&gt;We could specify certain subgroups and run separate regressions for each subgroup and obtain different estimates of $\tau$. To avoid false discoveries, &lt;mark&gt;this approach would require us to specify potential subgroups &lt;strong&gt;without&lt;/strong&gt; looking at the data.&lt;/mark&gt;
How can we use the data to inform us of potential subgroups?&lt;/p&gt;
&lt;p&gt;Let&amp;rsquo;s define,&lt;/p&gt;
&lt;p&gt;$$
Y_i=\textcolor{red}{ \tau\left(X_i\right) } W_i+f\left(X_i\right)+\varepsilon_i, \quad E\left[\varepsilon_i \mid X_i, W_i\right]=0,
$$&lt;/p&gt;
&lt;p&gt;where $\textcolor{red}{ \tau\left(X_i\right) }$ is the conditional ATE, i.e., $\textcolor{red}{ \tau\left(X_i\right) } :=\mathbb{E}\left[Y_i(1)-Y_i(0) \mid X_i\right]$. How do we estimate this?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Idea&lt;/strong&gt;: If we imagine we had access to some neighborhood $\mathcal{N}(x)$ where $\tau$ was constant, we could proceed exactly as before, by doing a residual-on-residual regression on the samples belonging to $\mathcal{N}(x)$, i.e.:&lt;/p&gt;
&lt;p&gt;$$
\tau(x) := \operatorname{lm}\left(Y_i - \hat{m}^{(-i)}(X_i) \sim W_i - \hat{e}^{(-i)}(X_i), \text{ weights } = \mathbf{1}\{X_i \in \mathcal{N}(x)\} \right)
$$&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;This is conceptually what Causal Forest does&lt;/strong&gt;, &lt;mark&gt;it estimates the treatment effect $\tau(x)$ for a target sample $X_i = x$ by running a weighted residual-on-residual regression on samples that have similar treatment effects.&lt;/mark&gt;&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    Recap: Causal Forest is running a a &amp;ldquo;forest&amp;rdquo;-localized version of Robinson’s regression.
  &lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;These &lt;em&gt;weights&lt;/em&gt; play a crucial role, so how does &lt;code&gt;grf&lt;/code&gt; 📦 find them?&lt;/p&gt;
&lt;h2 id=&#34;random-forest-as-an-adaptive-neighborhood-finder&#34;&gt;Random forest as an adaptive neighborhood finder&lt;/h2&gt;
&lt;p&gt;Breiman’s random forest for predicting the conditional mean $m(x) = \mathbb{E}(Y_i | X_i = x)$ can be briefly summarized in two steps:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Building phase: Build $B$ trees which greedily place covariate splits that &lt;mark&gt;maximize the squared difference in subgroups means&lt;/mark&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$$
n_L \cdot n_R \cdot ( \bar{y}_L - \bar{y}_R )^{2}
$$&lt;/p&gt;
&lt;ol start=&#34;2&#34;&gt;
&lt;li&gt;Prediction phase: Aggregate each tree’s prediction to form the final point estimate by averaging the outcomes $Y_i$ that fall into the same terminal leaf $L_b(X_i)$ as the targets sample $x$:&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;$$\begin{align}
\hat{m}(x) &amp;amp;= \frac{1}{B} \sum_{b=1}^B \sum_{i=1}^n   \frac{Y_i \mathbf{1}\left\{X_i \in L_b(x)\right\}}{\left|L_b(x)\right|} \\
&amp;amp;= \sum_{i=1}^n \frac{1}{B} \sum_{b=1}^B Y_i \frac{\mathbf{1}\left\{X_i \in L_b(x)\right\}}{\left|L_b(x)\right|} \tag{1} \\
&amp;amp;= \sum_{i=1}^n Y_i \textcolor{blue}{\frac{1}{B} \sum_{b=1}^B  \frac{\mathbf{1}\left\{X_i \in L_b(x)\right\}}{\left|L_b(x)\right|}} \\
&amp;amp; =\sum_{i=1}^n Y_i \textcolor{blue}{\alpha_i(x)} \tag{2},
\end{align}$$&lt;/p&gt;
&lt;p&gt;Note that, this procedure is a double summation, first over trees, then over training samples (see equation (1)). We can swap the order of summation and obtain $\textcolor{blue}{\alpha_i(x)}$ in the equation (2). &lt;mark&gt;We have defined $\textcolor{blue}{\alpha_i(x)}$ as the frequency with which the $i$-th training sample falls into the same leaf as $x$.&lt;/mark&gt;&lt;/p&gt;
&lt;p&gt;Does above $\textcolor{blue}{\alpha_i(x)}$ remind you the traditional deterministic kernel function and bandwidth story?&lt;/p&gt;
&lt;p&gt;The following image illustrates how the $\textcolor{blue}{\alpha_i(x)}$ are calculated: Some dots are larger because they are used by all trees, while some are smaller because they are only used by a few trees.&lt;/p&gt;
&lt;figure&gt;
    &lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/blogdown-image@main/uPic2/Screenshot%202024-08-20%20at%2020.38.30.png&#34; alt=&#34;weights from causal forests&#34; style=&#34;zoom:50%;&#34; /&gt;
    &lt;figcaption&gt;Figure 1: Weights from Causal Forests&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id=&#34;causal-forest&#34;&gt;Causal Forest&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Causal Forest&lt;/strong&gt; essentially combines Breiman (2001) and Robinson (1988) by modifying the steps above to:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;Building phase: Greedily places covariate splits that maximize the squared difference in subgroup treatment effects $$n_L \cdot n_R \cdot ( \hat{\tau}_L - \hat{\tau}_R )^{2},$$ where $\hat{\tau}$ is obtained by running Robinson’s residual-on-residual regression for each possible split point.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Use the resulting forest weights $\textcolor{blue}{\alpha_i(x)}$ to estimate $$\tau(x):=\operatorname{lm}\left(Y_i-\hat{m}^{(-i)}\left(X_i\right) \sim W_i-\hat{e}^{(-i)}\left(X_i\right), \text { weights }=\textcolor{blue}{\alpha_i(x)}\right),$$ where $\textcolor{blue}{\alpha_i(x)}$ capture how &amp;ldquo;similar&amp;rdquo; a target sample $x$ is to each of the training samples $X_i$&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;That is, &lt;mark&gt;&lt;em&gt;Causal Forest&lt;/em&gt; is running a “forest”-localized version of Robinson’s regression.&lt;/mark&gt; This adaptive weighting (instead of leaf-averaging) coupled with some other forest construction details known as &lt;strong&gt;“honesty” and “subsampling”&lt;/strong&gt; can be used to give asymptotic guarantees for estimation and inference with random forests (Wager &amp;amp; Athey, 2018)&lt;/p&gt;
&lt;h2 id=&#34;efficiently-estimating-summaries-of-the-cates&#34;&gt;Efficiently estimating summaries of the CATEs&lt;/h2&gt;
&lt;p&gt;What about estimating summaries of $\tau(x)$, in terms of estimands like the average treatment effect (ATE), or the best linear projection (BLP), that have guaranteed $\sqrt{n}$ rate of convergence along with exact confidence intervals?&lt;/p&gt;
&lt;p&gt;For estimating ATE, the most intuitive approach is to average the CATE, i.e. $\frac{1}{n}\sum_{i=1}^n \tau(X_i)$, right? However, there are more efficient methods than simply averaging individual CATE estimates.&lt;/p&gt;
&lt;p&gt;Robins, Rotnitzky &amp;amp; Zhao (1994) showed that the so-called &lt;mark&gt;&lt;strong&gt;Augmented Inverse Probability Weighted (AIPW) estimator is asymptotically optimal&lt;/strong&gt; for $\tau$ (meaning that among all non-parametric estimators, it has the lowest variance).&lt;/mark&gt;&lt;/p&gt;
$$\begin{gathered}
\hat{\tau}_{AIPW}=\frac{1}{n} \sum_{i=1}^n\left(\hat{\mu}_{(1)}\left(X_i\right)-\hat{\mu}_{(0)}\left(X_i\right)+\frac{W_i}{\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(1)}\left(X_i\right)\right)\right. \\ \left.-\frac{1-W_i}{1-\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(0)}\left(X_i\right)\right)\right), 
\end{gathered}$$
&lt;p&gt;where $$\mu_{(w)}(x):=\mu(x, w) := \mathbb{E}\left[Y_i \mid X_i=x, W_i=w\right]$$  and $$e(x)=\mathbb{P}\left[W_i=1 \mid X_i=x\right]$$&lt;/p&gt;
&lt;p&gt;To interpret the AIPW estimator $\hat{\tau}_{AIPW}$ , it is helpful to decompose it into two components: Let $\hat{\tau}_{AIPW} = A + B$ , where&lt;/p&gt;
$$
A=\frac{1}{n} \sum_{i=1}^n\left(\hat{\mu}_{(1)}\left(X_i\right)-\hat{\mu}_{(0)}\left(X_i\right)\right)
$$
$$
B=\frac{1}{n} \sum_{i=1}^n\left(\frac{W_i}{\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(1)}\left(X_i\right)\right)-\frac{1-W_i}{1-\hat{e}\left(X_i\right)}\left(Y_i-\hat{\mu}_{(0)}\left(X_i\right)\right)\right)
$$
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;$A$ represents the outcome regression adjustment estimator using $\hat{\mu}_{(w)}$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$B$ is an inverse propensity score weighting (IPW) estimator applied to the residuals $Y_i-\hat{\mu}_{\left(W_i\right)}\left(X_i\right)$&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The AIPW estimator utilizes propensity score weighting on the residuals to debias the direct estimate&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;One key property of the AIPW estimator is its &lt;strong&gt;&amp;ldquo;double robustness&amp;rdquo;&lt;/strong&gt;, which means that the estimator remains consistent and asymptotically normal even if either the outcome model or the propensity score model is misspecified. For proof, please refer to Stefan Wager&amp;rsquo;s Lecture 3 notes in “STATS 361: Causal Inference”.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The expression for above AIPW estimator can be rearranged and expressed as&lt;/p&gt;
$$
\begin{aligned}
&amp;\frac{1}{n} \sum_{i=1}^n\left(\tau\left(X_i\right) +
\textcolor{red}{\left[\frac{W_i}{e(X_i)} - \frac{1-W_i}{1-e(X_i)}\right]} 
\textcolor{violet}{\left[Y_i-\mu\left(X_i, W_i\right)\right]}

\right) \\

&amp; = \frac{1}{n} \sum_{i=1}^n\left(\tau\left(X_i\right)+\frac{W_i-e\left(X_i\right)}{e\left(X_i\right)\left[1-e\left(X_i\right)\right]}\left(Y_i-\mu\left(X_i, W_i\right)\right)\right) \\

&amp;\triangleq \frac{1}{n} \sum_{i=1}^n \Gamma_i 
\end{aligned}
$$
&lt;p&gt;We can understand above terms as:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;$\tau(X_i)$ is an initial treatment effect estimate&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\textcolor{red}{\left[\frac{W_i}{e(X_i)} - \frac{1-W_i}{1-e(X_i)}\right]}$ is the &lt;strong&gt;Riesz Representer&lt;/strong&gt; (Chernozhukov et al., 2022), which is used to &amp;ldquo;correct&amp;rdquo; the bias&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\textcolor{violet}{\left[Y_i-\mu\left(X_i, W_i\right)\right]}$ is the residual part&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;$\Gamma_i$ is called double robust score&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Chernozhukov, Victor, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. 2018. “Double/Debiased Machine Learning for Treatment and Structural Parameters.” &lt;i&gt;The Econometrics Journal&lt;/i&gt; 21 (1): C1–68. &lt;a href=&#34;https://doi.org/10.1111/ectj.12097&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1111/ectj.12097&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Athey, Susan, Julie Tibshirani, and Stefan Wager. 2019. “Generalized Random Forests.” &lt;i&gt;The Annals of Statistics&lt;/i&gt; 47 (2): 1148–78. &lt;a href=&#34;https://doi.org/10.1214/18-AOS1709&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1214/18-AOS1709&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Robinson, Peter M. “Root-N-consistent semiparametric regression.” Econometrica: Journal of the Econometric Society (1988): 931-954.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Wager, Stefan, and Susan Athey. “Estimation and inference of heterogeneous treatment effects using random forests.” Journal of the American Statistical Association 113.523 (2018): 1228-1242.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Wager, S. (2022). STATS 361: Causal Inference Lecture notes. Stanford University. &lt;a href=&#34;https://web.stanford.edu/~swager/stats361.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://web.stanford.edu/~swager/stats361.pdf&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Chernozhukov, V., Newey, W. K., &amp;amp; Singh, R. (2022). Automatic Debiased Machine Learning of Causal and Structural Effects. &lt;i&gt;Econometrica&lt;/i&gt;, &lt;i&gt;90&lt;/i&gt;(3), 967–1027. &lt;a href=&#34;https://doi.org/10.3982/ECTA18515&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.3982/ECTA18515&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>From Donsker Classes to Neyman Orthogonality: The Power of DML</title>
      <link>https://chenxing.space/blog/from-donsker-classes-to-neyman-orthogonality-the-power-of-dml/</link>
      <pubDate>Fri, 26 Apr 2024 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/from-donsker-classes-to-neyman-orthogonality-the-power-of-dml/</guid>
      <description>&lt;h3 id=&#34;motivation--intuition&#34;&gt;Motivation &amp;amp; Intuition&lt;/h3&gt;
&lt;p&gt;In classical semiparametric theory, we want to estimate a low‑dimensional target parameter (say, a treatment effect) while controlling for high‑dimensional nuisance functions (like nonparametric regressions). However, in order to use the central limit theorem (CLT) and to characterize the asymptotic behavior, classical results require that the space of functions in which these nuisance functions lie is “small” in a technical sense. &lt;mark&gt;In particular, they must form a &lt;strong&gt;Donsker class&lt;/strong&gt; — roughly speaking, a collection of functions whose complexity (measured via “entropy”) is bounded enough so that the empirical process converges to a Gaussian process. This condition, however, is too restrictive in modern applications where the nuisance functions are estimated by flexible machine learning methods&lt;/mark&gt; (e.g., random forests, boosting, deep neural nets) that may come from very large, high‑dimensional spaces.&lt;/p&gt;
&lt;p&gt;The double machine learning approach overcomes this problem by using &lt;strong&gt;Neyman orthogonal&lt;/strong&gt; scores. The key idea is that the moment functions used to estimate the target parameter are constructed in such a way that &lt;mark&gt;small errors in estimating the nuisance functions have only a second-order effect on the final estimator&lt;/mark&gt;. In plain language, even if your machine learning methods are “messy” or come from huge function classes (i.e. they do not satisfy the Donsker conditions), the estimation of your target parameter remains robust as long as the nuisance estimators converge at a certain rate. Moreover, the approach uses sample splitting (or cross‑fitting) to avoid overfitting biases.&lt;/p&gt;
&lt;h3 id=&#34;key-math-details&#34;&gt;Key Math Details&lt;/h3&gt;
&lt;center&gt;&lt;img src=&#34;https://cdn.jsdelivr.net/gh/chenx2018/cloudimg@main/uPic/image-20250220090505509.png&#34; alt=&#34;image-20250220090505509&#34; style=&#34;zoom:100%;&#34; /&gt;&lt;/center&gt;
&lt;h3 id=&#34;summary&#34;&gt;Summary&lt;/h3&gt;
&lt;p&gt;Flexibility in Nuisance Estimation: DML frees us from the need for nuisance estimators to lie in “small” Donsker classes. Thanks to Neyman orthogonality and cross-fitting, we can plug in flexible, machine-learning based nuisance estimates—even if they come from very rich function spaces—without contaminating the asymptotic distribution of the target parameter estimator.&lt;/p&gt;
&lt;p&gt;Thus, the contribution of DML is in allowing the use of complex, modern ML methods to estimate nuisance functions while still obtaining valid inference for the target parameter, bypassing the traditional, more restrictive Donsker conditions.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Learning Resource: Causal Inference</title>
      <link>https://chenxing.space/blog/2023-12-14-learning-resource-causal-inference/learning-resource-causal-inference/</link>
      <pubDate>Thu, 14 Dec 2023 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/2023-12-14-learning-resource-causal-inference/learning-resource-causal-inference/</guid>
      <description>&lt;h2 id=&#34;difference-between-econometrics--statistics&#34;&gt;Difference between Econometrics &amp;amp; Statistics&lt;/h2&gt;
&lt;p&gt;What is the difference between Econometrics and Statistics? Professor Joshua Angrist (MIT) explains the difference in this &lt;a href=&#34;https://www.youtube.com/watch?v=uVrr_-UUgWk&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;video&lt;/a&gt;.&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    Statisticians use sampling to make &lt;strong&gt;statistical inferences&lt;/strong&gt; about large populations. Econometricians, on the other hand, examine counterfactuals to make &lt;strong&gt;causal inferences&lt;/strong&gt;.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h2 id=&#34;causal-inference-theory-learning-resources&#34;&gt;Causal Inference (theory) learning resources&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Golub Capital Social Impact Lab (2023). Machine Learning-based Causal Inference Tutorial.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://bookdown.org/stanfordgsbsilab/ml-ci-tutorial/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;textbook&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.gsb.stanford.edu/faculty-research/centers-initiatives/sil/research/methods/ai-machine-learning/short-course&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;YouTube tutorial&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Key idea to remember
&lt;ul&gt;
&lt;li&gt;FWL Theorem ➜ Robinson&amp;rsquo;s Transformation ➜ R-learner (with R loss)&lt;/li&gt;
&lt;li&gt;See &lt;a href=&#34;https://web.stanford.edu/~swager/stats361.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;STATS 361: Causal Inference&lt;/a&gt;, page 36 for details. 















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/uPic/Screenshot%202024-01-20%20at%2008.10.47-20240120082012928.png&#34; alt=&#34;Screenshot 2024-01-20 at 08.10.47&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/li&gt;
&lt;li&gt;For R-learner, &lt;u&gt;Nie, Xinkun, and Stefan Wager. Quasi-oracle estimation of heterogeneous treatment effects. Biometrika&lt;/u&gt; provides a good summary.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Peng Ding&amp;rsquo;s textbook &lt;a href=&#34;https://arxiv.org/abs/2305.18793&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;&lt;strong&gt;A first course in causal inference&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Wager, S. (2022). &lt;a href=&#34;https://web.stanford.edu/~swager/stats361.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;STATS 361: Causal Inference&lt;/a&gt;. Lecture notes, Stanford University.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.masteringmetrics.com/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Mastering &amp;lsquo;Metrics&lt;/a&gt;, less theoretical.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://matheusfacure.github.io/python-causality-handbook/landing-page.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Causal Inference for The Brave and True&lt;/a&gt; is an open-source resource primarily focused on econometrics and the statistics of science. &lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www2.stat.duke.edu/~fl35/CausalInferenceClass.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Causal Inference - Statistical Science - Duke University&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;The slides are provided by Professor Fan Li.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;I highly recommend to read these slides as summary to get a big picture and read Peng Ding&amp;rsquo;s textbook for details and rigorous proofs.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.bradyneal.com/causal-inference-course&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Introduction to Causal Inference (Fall 2020) by Brady Neal&lt;/a&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Brady kindly provides all his course material and &lt;a href=&#34;https://www.youtube.com/playlist?list=PLoazKTcS0Rzb6bb9L508cyJ1z-U9iWkA0&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;YouTube tutorials&lt;/a&gt;. They are super helfpul!!!&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;causal-inference-in-r&#34;&gt;Causal Inference in R&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;R 📦 &lt;a href=&#34;https://cran.r-project.org/web/views/CausalInference.html#ate&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;CRAN Task View: Causal Inference&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Propensity Score Weighting tutorial: &lt;a href=&#34;https://www.andrewheiss.com/blog/2020/12/01/ipw-binary-continuous/#ipw-with-the-ipw-package-binary-treatment&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Generating inverse probability weights for both binary and continuous treatments&lt;/a&gt;. In this tutorial, the author introduces &lt;a href=&#34;https://cran.r-project.org/package=ipw&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;&lt;strong&gt;ipw&lt;/strong&gt;&lt;/a&gt; and &lt;a href=&#34;https://github.com/ngreifer/WeightIt&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;&lt;strong&gt;WeightIt&lt;/strong&gt;&lt;/a&gt; 📦s.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/xnie/rlearner&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;rlearner&lt;/a&gt; for Quasi-Oracle Estimation of Heterogeneous Treatment Effects.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://docs.doubleml.org/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;DoubleML &amp;mdash; DoubleML documentation&lt;/a&gt; R 📦. Paper: &lt;a href=&#34;https://arxiv.org/pdf/2103.09603.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;DoubleML - An Object-Oriented Implementation of Double Machine Learning in R&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://www.r-causal.org/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Causal Inference in R&lt;/a&gt; is a bookdown tutorial. It is new and incomplete.&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://grf-labs.github.io/grf/index.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;grf&lt;/a&gt; package for generalized random forests.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;causal-inference-in-python&#34;&gt;Causal Inference in Python&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;🐍 &lt;a href=&#34;https://econml.azurewebsites.net/spec/spec.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;EconML User Guide&lt;/a&gt;, this is for double machine learning.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;difference-in-difference&#34;&gt;Difference in Difference&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Here is the shared Dropbox from Professor Jeffrey Wooldridge: &lt;a href=&#34;https://www.dropbox.com/scl/fo/xvuiqj1g910tom6tlyy5g/ADxMDZPO7WCmN22HLjnLJL8?rlkey=liydy17cs6y1s24hzr2ride3v&amp;amp;e=1&amp;amp;st=7mrqfy0d&amp;amp;dl=0&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;share_jeff_wooldridge_did&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Here is a link to a list of Stata packages that cover the recent literature on staggered DiD designs: &lt;a href=&#34;https://asjadnaqvi.github.io/DiD/docs/01_stata/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://asjadnaqvi.github.io/DiD/docs/01_stata/&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;YouTube Tutorial: &lt;a href=&#34;https://www.youtube.com/watch?v=q7fpkYcUu1g&amp;amp;t=17s&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Jeff Wooldridge presents &amp;ldquo;Differences in Differences&amp;rdquo; to the ASA Ann Arbor Chapter&lt;/a&gt;. Professor Wooldridge provides us a &lt;strong&gt;clever transformation&lt;/strong&gt; that enable us to convert the panel data to cross-sectional data, then one can apply their favorite treatment effect estimators such as matching, IPW, AIPW and etc. Specifically,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Identical to regression adjustment (RA) on a transformed variable:&lt;/li&gt;
&lt;/ul&gt;

  $$
  \dot{Y}_{i t}=Y_{i t}-\frac{1}{(q-1)} \sum_{s=1}^{q-1} Y_{i s}, \quad t \geq q
  $$
  
&lt;ul&gt;
&lt;li&gt;For any treatment period $t \geq q$, apply standard treatment effect methods to&lt;/li&gt;
&lt;/ul&gt;
 
  $$
  \left\{\left(\dot{Y}_{i t}, D_i, \mathbf{X}_i\right): i=1, \ldots, N\right\}
  $$
  
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Inverse probability weighting&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;IPWRA (Doubly Robust)&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Covariate or PS matching&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For more details, please check the paper: &lt;a href=&#34;https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4516518&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;A Simple Transformation Approach to Difference-in-Differences Estimation for Panel Data&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;synthetic-control&#34;&gt;Synthetic Control&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;YouTube tutorial: &lt;a href=&#34;https://www.youtube.com/watch?v=xCNQdnZzg64&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;How To Use The Synthetic Control Method in R Step-By-Step: Effect of California&amp;rsquo;s Tobacco Program&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;YouTube tutorial: &lt;a href=&#34;https://www.youtube.com/watch?v=EIBR0kpTrfE&amp;amp;ab_channel=KosukeImai&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Synthetic Control Method&lt;/a&gt;. This tutorial is short but provides the key insights.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://carlos-mendez.quarto.pub/r-synthetic-control-tutorial/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://carlos-mendez.quarto.pub/r-synthetic-control-tutorial/&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Good summary for &lt;a href=&#34;https://bookdown.org/mike/data_analysis/synthetic-difference-in-differences.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;synthetic did (sdid)&lt;/a&gt;. In this post, it compares the DID, SC and SDID&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>Learning Resource: Causal Machine Learning with DoubleML</title>
      <link>https://chenxing.space/blog/learning-resource-causal-machine-learning-with-doubleml/</link>
      <pubDate>Tue, 21 Nov 2023 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/learning-resource-causal-machine-learning-with-doubleml/</guid>
      <description>&lt;script src=&#34;https://chenxing.space/blog/learning-resource-causal-machine-learning-with-doubleml/index.en_files/fitvids/fitvids.min.js&#34;&gt;&lt;/script&gt;
&lt;p&gt;Here are some study notes for &lt;strong&gt;Double Machine Learning in causal inference&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;Double machine learning, as introduced by (Chernozhukov et al. 2018), is a methodology used in causal inference, which is particularly useful when dealing with high-dimensional data.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;The Challenge in High-Dimensional Data&lt;/strong&gt;: In causal inference, we often want to estimate the effect of a particular variable (the treatment) on an outcome. However, in high-dimensional settings, where we have a large number of potential control variables (features), traditional methods can be inefficient or biased.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Concept of Double Machine Learning&lt;/strong&gt;: The term “double” refers to a two-step process. The first step is to use machine learning methods to predict both the treatment and the outcome based on the control variables. The second step is to use these predictions to correct the bias in estimating the causal effect.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;First Step – Machine Learning Models&lt;/strong&gt;: Here, we use machine learning algorithms to model the relationship between the control variables and (a) the treatment, and (b) the outcome. This helps in understanding how these control variables influence both the treatment and the outcome separately.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Second Step – Causal Effect Estimation&lt;/strong&gt;: After we’ve modeled the treatment and outcome separately, we can now adjust our estimation of the causal effect. This adjustment is crucial because it accounts for the influence of the control variables, reducing the risk of bias that could be introduced if these variables were ignored or improperly handled.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Orthogonality Principle&lt;/strong&gt;: A key aspect of double machine learning is the orthogonality principle. It ensures that the estimation of the causal effect is not heavily influenced by small changes in the estimation of the control variables. This makes the method robust to the choice of machine learning models used in the first step.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Advantage in Real-world Applications&lt;/strong&gt;: In practical scenarios, especially in marketing and economics, where datasets are large and complex, double machine learning offers a more reliable way to discern causal relationships, as it efficiently handles a large number of variables and complex relationships between them.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;In summary, double machine learning in causal inference provides a robust and efficient way to estimate causal effects in high-dimensional settings by leveraging the power of machine learning algorithms to control for a large set of variables, thus reducing bias and increasing the reliability of causal estimations.&lt;/p&gt;
&lt;h2 id=&#34;learning-resource&#34;&gt;Learning Resource&lt;/h2&gt;
&lt;p&gt;&lt;a href=&#34;https://docs.doubleml.org/tutorial/stable/2023/09/25/welcome-tools-for-causality/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;This website&lt;/a&gt; for the “Tools for Causality - Double Machine Learning” course offers materials on Double Machine Learning, causal machine learning, heterogeneous treatment effects, sensitivity analysis, and advanced methods like Difference-in-Differences.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;There is also a helpful &lt;strong&gt;tutorial on YouTube&lt;/strong&gt;, check &lt;a href=&#34;https://www.youtube.com/watch?v=1bHzi6ucnr0&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Some useful slides they provided are the following:&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;div class=&#34;shareagain&#34; style=&#34;min-width:300px;margin:1em auto;&#34; data-exeternal=&#34;1&#34;&gt;
&lt;iframe src=&#34;https://docs.doubleml.org/tutorial/stable/slides/part1/Lect1_Introduction_to_DML.html#1&#34; width=&#34;1600&#34; height=&#34;900&#34; style=&#34;border:2px solid currentColor;&#34; loading=&#34;lazy&#34; allowfullscreen&gt;&lt;/iframe&gt;
&lt;script&gt;fitvids(&#39;.shareagain&#39;, {players: &#39;iframe&#39;});&lt;/script&gt;
&lt;/div&gt;
&lt;p&gt;Another similar Good Tutorial: &lt;a href=&#34;https://docs.doubleml.org/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;DoubleML — DoubleML documentation&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;div id=&#34;refs&#34; class=&#34;references csl-bib-body hanging-indent&#34; entry-spacing=&#34;0&#34;&gt;
&lt;div id=&#34;ref-ChernozhukovUnknownTitle2018&#34; class=&#34;csl-entry&#34;&gt;
&lt;p&gt;Chernozhukov, Victor, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. 2018. “Double/Debiased Machine Learning for Treatment and Structural Parameters.” &lt;em&gt;The Econometrics Journal&lt;/em&gt;. &lt;a href=&#34;https://doi.org/10.1111/ectj.12097&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;https://doi.org/10.1111/ectj.12097&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
</description>
    </item>
    
  </channel>
</rss>
