<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>distributions | Chen Xing</title>
    <link>https://chenxing.space/tag/distributions/</link>
      <atom:link href="https://chenxing.space/tag/distributions/index.xml" rel="self" type="application/rss+xml" />
    <description>distributions</description>
    <generator>Wowchemy (https://wowchemy.com)</generator><language>en-us</language><lastBuildDate>Wed, 03 Aug 2022 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://chenxing.space/media/sharing.png</url>
      <title>distributions</title>
      <link>https://chenxing.space/tag/distributions/</link>
    </image>
    
    <item>
      <title>Beta Distribution — Intuition, Derivation, and Examples</title>
      <link>https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/</link>
      <pubDate>Wed, 03 Aug 2022 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/</guid>
      <description>&lt;h2 id=&#34;motivation&#34;&gt;Motivation&lt;/h2&gt;
&lt;h4 id=&#34;model-probabilities&#34;&gt;Model probabilities&lt;/h4&gt;
&lt;p&gt;The Beta distribution is &lt;strong&gt;a probability distribution on probabilities&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The Beta distribution can be understood as representing a distribution &lt;em&gt;of probabilities&lt;/em&gt;, that is, it represents all the possible values of a probability when we don&amp;rsquo;t know what that probability is. For example,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;the Click-Through Rate of your advertisement&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;the conversion rate of customers actually purchasing in your store&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;how likely the customer will become &amp;ldquo;inactive&amp;rdquo;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Because the Beta distribution models a probability, its domain is bounded between &lt;strong&gt;0&lt;/strong&gt; and &lt;strong&gt;1&lt;/strong&gt;.&lt;/p&gt;
&lt;h4 id=&#34;generalization-of-uniform-distribution&#34;&gt;Generalization of Uniform Distribution&lt;/h4&gt;
&lt;p&gt;Give me a &lt;strong&gt;continuous&lt;/strong&gt; and &lt;strong&gt;bounded&lt;/strong&gt; random variable except the &lt;em&gt;Uniform Distribution&lt;/em&gt;. This is another way to look at &lt;em&gt;beta distribution&lt;/em&gt;, continuous and bounded between 0 and 1; also the density is not flat.&lt;/p&gt;

$$
X \sim Beta(a, b), \text{ where } a&gt;0, \ b&gt;0.
$$



$$
f_X(x) = c \cdot x ^{a-1}(1-x)^{b-1}, \text{ where } x&gt;0.
$$


&lt;p&gt;What is $c$ ? Just a normalization constant! We&amp;rsquo;ll find the value of $c$ later.&lt;/p&gt;
&lt;h4 id=&#34;conjugate-prior&#34;&gt;Conjugate Prior&lt;/h4&gt;
&lt;p&gt;The Beta distribution is the &lt;strong&gt;conjugate prior&lt;/strong&gt; for the Bernoulli, binomial, negative binomial and geometric distributions (seems like those are the distributions that involve success &amp;amp; failure) in Bayesian inference.&lt;/p&gt;
&lt;mark&gt;Computing a posterior using a conjugate prior is very convenient, because you can avoid expensive numerical computation involved in Bayesian Inference.&lt;/mark&gt;
&lt;blockquote&gt;
&lt;p&gt;Conjugate prior = Convenient prior&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;For example, the beta distribution is a conjugate prior to the binomial. &lt;strong&gt;If we choose to use the beta distribution Beta(α, β) as a prior, during the modeling phase, we already know the posterior will also be a beta distribution.&lt;/strong&gt; Therefore, after carrying out more experiments, &lt;strong&gt;you can compute the posterior simply by adding the number of successes (x), and failures (n-x) to the existing parameters α, β respectively&lt;/strong&gt;, instead of multiplying the likelihood with the prior distribution. The posterior also becomes a Beta distribution with parameters &lt;strong&gt;(x+α, n-x+β).&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id=&#34;what-is-the-intuition&#34;&gt;What is the Intuition?&lt;/h2&gt;
&lt;p&gt;The intuition for the beta distribution comes into play when we look at it from the lens of the binomial distribution.&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/1*eZKUz1_Jyvt6DNj8Tcv20Q.png&#34; alt=&#34;img&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;The difference between the binomial and the beta is that the &lt;strong&gt;former models the number of successes (x), while the latter models the probability (p) of success.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;In other words, the probability is a &lt;strong&gt;parameter&lt;/strong&gt; in binomial; In the Beta, the probability is a &lt;strong&gt;random variable&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id=&#34;interpretation-of-α-β&#34;&gt;Interpretation of &lt;strong&gt;α, β&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;You can think of &lt;strong&gt;α-1 as the number of successes&lt;/strong&gt; and &lt;strong&gt;β-1 as the number of failures,&lt;/strong&gt; just like &lt;strong&gt;n&lt;/strong&gt; &amp;amp; &lt;strong&gt;n-x&lt;/strong&gt; terms in binomial.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;You can choose the α and β parameters however you think they are supposed to be&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If you think the probability of success is very high, let’s say 90%, &lt;strong&gt;set 90 for α&lt;/strong&gt; and &lt;strong&gt;10 for β.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;If you think otherwise, 90 for β and 10 for α.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;As &lt;strong&gt;α&lt;/strong&gt; becomes larger (more successful events), the bulk of the probability distribution will shift towards the right, whereas an increase in &lt;strong&gt;β&lt;/strong&gt; moves the distribution towards the left (more failures).&lt;/p&gt;
&lt;p&gt;Also, the distribution will narrow if both &lt;strong&gt;α&lt;/strong&gt; and &lt;strong&gt;β&lt;/strong&gt; increase, for we are more certain.&lt;/p&gt;
&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/%E6%88%AA%E5%B1%8F2022-08-03%2010.48.34.png&#34; alt=&#34;截屏2022-08-03 10.48.34&#34; style=&#34;zoom:50%;&#34; /&gt;
&lt;p&gt;Dr. Bognar at the University of Iowa built &lt;a href=&#34;https://homepage.divms.uiowa.edu/~mbognar/applets/beta.html&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;the calculator for Beta distribution&lt;/a&gt;, which I found useful and beautiful. You can experiment with different values of &lt;strong&gt;α&lt;/strong&gt; and &lt;strong&gt;β&lt;/strong&gt; and visualize how the shape changes.&lt;/p&gt;
&lt;h2 id=&#34;derivation&#34;&gt;Derivation&lt;/h2&gt;
&lt;p&gt;In this section, we&amp;rsquo;ll derive Beta distribution using the Beta-Gamma Connections.&lt;/p&gt;
&lt;div class=&#34;alert alert-tip&#34;&gt;
  &lt;div&gt;
    We can regard Beta distribution as the &lt;strong&gt;fraction of waiting time&lt;/strong&gt;.
  &lt;/div&gt;
&lt;/div&gt;
&lt;h3 id=&#34;fraction-of-waiting-time&#34;&gt;Fraction of Waiting Time&lt;/h3&gt;
&lt;p&gt;Let $X$  be the waiting time at Bank,&lt;/p&gt;

$$
X \sim Gamma(n_1, \lambda)
$$

&lt;p&gt;Let $Y$  be the waiting time at Post Office,&lt;/p&gt;

$$
Y \sim Gamma(n_2, \lambda)
$$

&lt;p&gt;Assume $X$  and $Y$  are independent. &lt;strong&gt;What is the distribution of the proportion&lt;/strong&gt; $\frac{X}{X+Y}$ ?&lt;/p&gt;
&lt;p&gt;Solution:&lt;/p&gt;
&lt;p&gt;Let $T := X+Y$ be the total waiting time. Clearly, $T \sim Gamma(n_1+n_2, \lambda)$, you can prove it by MGF.&lt;/p&gt;
&lt;p&gt;Let $W =: \frac{X}{X+Y}$ be the proportion of waiting time at Bank to the total waiting time. We need to find the PDF of $W$.&lt;/p&gt;
&lt;p&gt;The idea is to find the joint PDF $f_{T,W}(t,w)$ at first, and then get the marginal distribution.&lt;/p&gt;

$$
\begin{aligned}f_{T,W}(t,w) &amp;= f_{X,Y}(x,y) \left | \frac{\partial(x,y)}{\partial(t,w)} \right|\\
    &amp;= \frac{1}{\Gamma(n_1)}\lambda^{n_1}x^{n_1 - 1}e^{-\lambda x} \frac{1}{\Gamma(n_2)}\lambda^{n_2}x^{n_2 - 1}e^{-\lambda y} \left|-t\right|\\
    &amp;= \lambda^{n_1+n_2}t^{n_1+n_2-1}e^{-\lambda t} \frac{1}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}\\
    &amp;= \frac{\lambda^{n_1+n_2}t^{n_1+n_2-1}e^{-\lambda t}}{\Gamma(n_1+n_2)} \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}\\
    &amp;= f_T(t) \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}\end{aligned}
$$

&lt;p&gt;Integrating $t$ out to get the marginal:&lt;/p&gt;

$$
\begin{aligned}f_W(w) &amp;= \int_0^\infty f_{T,W}(t,w) dt \\&amp;= \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} \cdot\int_0^\infty f_T(t)dt \\&amp;= \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} \end{aligned}
$$

&lt;p&gt;Here we have Beta,&lt;/p&gt;
$$W \sim Beta(n_1, n_2)$$
&lt;mark&gt;REMARK: the above result also proves &lt;strong&gt;W and T and independent&lt;/strong&gt;!&lt;/mark&gt;
&lt;h3 id=&#34;beta-function-as-a-normalizing-constant&#34;&gt;Beta Function as a normalizing constant&lt;/h3&gt;
&lt;p&gt;Note that, $f_W(w)$  is a PDF needed to be integrated to 1,&lt;/p&gt;

$$
\int_0^1\frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} dw \equiv 1
$$

&lt;p&gt;So the normalization constant should be,&lt;/p&gt;

$$
c = \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)} := \frac{1}{B(n_1, n_2)}
$$

&lt;h3 id=&#34;mean-of-beta-distribution&#34;&gt;Mean of Beta distribution&lt;/h3&gt;
&lt;p&gt;As a byproduct in the above derivation, we get that fact that &lt;strong&gt;W and T are independent&lt;/strong&gt;. Then, we can use this to derive the mean of Beta distribution,&lt;/p&gt;
$$E(WT) = E(W)E(T) \implies E(X) =E\left(\frac{X}{X+Y}\right)E(X+Y),$$
&lt;p&gt;rearrange to get,&lt;/p&gt;
$$E\left(\frac{X}{X+Y}\right) = \frac{E(X)}{E(X+Y)}$$
&lt;p&gt;This result is clear Not True in general, but under our setting, we have this interesting result.&lt;/p&gt;
&lt;p&gt;We can use this result to find the mean of $W \sim Beta(a, b)$ without the slightest trace of calculus.&lt;/p&gt;
$$E(W)=E\left(\frac{X}{X+Y}\right)=\frac{E(X)}{E(X+Y)}=\frac{a / \lambda}{a / \lambda+b / \lambda}=\frac{a}{a+b}$$
&lt;h3 id=&#34;getting-beta-parameters-in-practice&#34;&gt;Getting Beta parameters in practice&lt;/h3&gt;
&lt;p&gt;The Beta distribution is the conjugate prior for many common distributions. We use it a lot. But in practice, &lt;mark&gt;how to figure out its parameters?&lt;/mark&gt; It is sometimes useful to estimate quickly the parameters of the Beta distribution using the method of moments:&lt;/p&gt;
$$X \sim Beta(\alpha, \beta),$$
$$\alpha + \beta = \frac{E(X)(1-E(X))}{Var(X)} - 1,$$
$$\alpha = (\alpha + \beta)E(X),$$
$$\beta = (\alpha + \beta)(1 - E(X))$$
&lt;p&gt;Here is the R code:&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# calculate beta params using method of moments&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;cal_beta_params&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;function&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;sdX&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;p&#34;&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;n&#34;&gt;varX&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;sdX^2&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;n&#34;&gt;sum_ab&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;*&lt;/span&gt; &lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;m&#34;&gt;1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;/&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;varX&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;-&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;n&#34;&gt;a&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;sum_ab&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;*&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;n&#34;&gt;b&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;sum_ab&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;*&lt;/span&gt; &lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;m&#34;&gt;1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;c1&#34;&gt;# return&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;c&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;s&#34;&gt;&amp;#34;shape&amp;#34;&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;a&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;scale&amp;#34;&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;b&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;p&#34;&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# example&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;cal_beta_params&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;meanX&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;0.136&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;sdX&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;0.103&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;##    shape    scale 
## 1.370320 8.705559
&lt;/code&gt;&lt;/pre&gt;&lt;h3 id=&#34;summary&#34;&gt;Summary&lt;/h3&gt;
&lt;p&gt;To summarize, the bank–post office story tells us that: when we add independent Gamma r.v.s $X$  and $Y$  with the same rate $\lambda$ ,&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;the total $X+Y$ has a Gamma distribution;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;the fraction $\frac{X}{X+Y}$ has a Beta distribution;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;the total is independent of the fraction.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;examples&#34;&gt;Examples&lt;/h2&gt;
&lt;p&gt;The PDF of Beta distribution can be U-shaped with asymptotic ends, bell-shaped, strictly increasing/decreasing or even straight lines. As you change &lt;strong&gt;α&lt;/strong&gt; or &lt;strong&gt;β&lt;/strong&gt;, the shape of the distribution changes.&lt;/p&gt;
&lt;h3 id=&#34;i-bell-shape&#34;&gt;I. Bell-Shape&lt;/h3&gt;
&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/index.en_files/figure-html/bellshape-1.png&#34; width=&#34;672&#34; /&gt;
&lt;p&gt;The PDF of a beta distribution is approximately normal if &lt;strong&gt;α&lt;/strong&gt; + &lt;strong&gt;β&lt;/strong&gt; is large enough and α &amp;amp; β are approximately equal.&lt;/p&gt;
&lt;h4 id=&#34;intuition-behind-bell-shape&#34;&gt;Intuition behind Bell-Shape&lt;/h4&gt;
&lt;p&gt;&lt;strong&gt;Why would Beta(2,2) be bell-shaped?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;If you think:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;α-1 as the number of successes&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;β-1 as the number of failures&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Beta(2,2)&lt;/strong&gt; means you got 1 success and 1 failure&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;So it makes sense that the probability of the success is highest at 0.5.&lt;/p&gt;
&lt;p&gt;Also, &lt;strong&gt;Beta(1,1)&lt;/strong&gt; would mean you got zero for the head and zero for the tail. Then, your guess about the probability of success should be the same throughout [0,1]. The horizontal straight line confirms it.&lt;/p&gt;
&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/index.en_files/figure-html/beta11-1.png&#34; width=&#34;672&#34; /&gt;
&lt;h3 id=&#34;ii-straight-lines&#34;&gt;II. Straight Lines&lt;/h3&gt;
&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/index.en_files/figure-html/unnamed-chunk-2-1.png&#34; width=&#34;672&#34; /&gt;
&lt;p&gt;&lt;strong&gt;α = 1 or β = 1&lt;/strong&gt;, the beta PDF can be a straight line.&lt;/p&gt;
&lt;h3 id=&#34;iii-u-shape&#34;&gt;III. U-Shape&lt;/h3&gt;
&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-intuition-derivation-and-examples/index.en_files/figure-html/unnamed-chunk-3-1.png&#34; width=&#34;672&#34; /&gt;
When $\alpha &lt; 1, \beta&lt;1$ the PDF of the Beta is U-shaped.
&lt;h2 id=&#34;reference&#34;&gt;Reference&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://en.wikipedia.org/wiki/Beta_distribution&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;wiki Beta distribution&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://towardsdatascience.com/beta-distribution-intuition-examples-and-derivation-cf00f4db57af&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Beta Distribution — Intuition, Examples, and Derivation&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://stats.stackexchange.com/questions/47771/what-is-the-intuition-behind-beta-distribution&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;What is the intuition behind beta distribution?&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=UZjlBQbV1KU&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;youtube video Lecture 23: Beta distribution | Statistics 110&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=v1uUgTcInQk&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;another youtube video: Beta distribution - an introduction&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
</description>
    </item>
    
    <item>
      <title>Beta distribution as a prior</title>
      <link>https://chenxing.space/blog/beta-distribution-as-a-prior/</link>
      <pubDate>Mon, 25 Jul 2022 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/beta-distribution-as-a-prior/</guid>
      <description>&lt;h2 id=&#34;introduction&#34;&gt;Introduction&lt;/h2&gt;
&lt;p&gt;The Beta distribution is a useful probability distribution when you want model uncertainty over a parameter bounded between 0 and 1.&lt;/p&gt;
&lt;p&gt;In this post, we&amp;rsquo;ll explore how the two parameters of the Beta distribution determine its shape.&lt;/p&gt;
&lt;p&gt;One way to see how the shape parameters of the Beta distribution affect its shape is to generate a large number of random draws using the &lt;code&gt;rbeta(n, shape1, shape2)&lt;/code&gt; function and visualize these as a histogram.&lt;/p&gt;
&lt;h2 id=&#34;beta11&#34;&gt;Beta(1,1)&lt;/h2&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Explore using the rbeta function&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;rbeta&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;n&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1000000&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape2&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Visualize the results&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;hist&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-as-a-prior/index.en_files/figure-html/unnamed-chunk-1-1.png&#34; width=&#34;672&#34; /&gt;
&lt;p&gt;A Beta(1,1) distribution is the same as a uniform distribution between 0 and 1. It is useful as a so-called &lt;em&gt;non-informative&lt;/em&gt; prior as it expresses than any value from 0 to 1 is equally likely.&lt;/p&gt;
&lt;h2 id=&#34;discover-shape-parameters&#34;&gt;Discover shape parameters&lt;/h2&gt;
&lt;p&gt;What happened if you set shape parameters to negative numbers?&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Modify the parameters&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;rbeta&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;n&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1000000&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;-1&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape2&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Explore the results&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;head&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;## [1] NaN NaN NaN NaN NaN NaN
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Yes, &lt;code&gt;NaN&lt;/code&gt; stands for &lt;em&gt;not a number&lt;/em&gt; and the reason you got a lot of &lt;code&gt;NaN&lt;/code&gt;s is that the Beta distribution is only defined when its shape parameters are &lt;strong&gt;positive&lt;/strong&gt;.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Modify the parameters&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;rbeta&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;n&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1000000&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;100&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape2&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;100&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Visualize the results&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;hist&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-as-a-prior/index.en_files/figure-html/unnamed-chunk-3-1.png&#34; width=&#34;672&#34; /&gt;
&lt;p&gt;So the larger the shape parameters are, the more concentrated the beta distribution becomes. When used as a prior, this Beta distribution encodes the information that the parameter is most likely close to 0.5 .&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Modify the parameters&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;rbeta&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;n&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1000000&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;100&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;shape2&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;20&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# Visualize the results&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;hist&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;beta_sample&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;img src=&#34;https://chenxing.space/blog/beta-distribution-as-a-prior/index.en_files/figure-html/unnamed-chunk-4-1.png&#34; width=&#34;672&#34; /&gt;
&lt;p&gt;So the larger the &lt;code&gt;shape1&lt;/code&gt; parameter is the closer the resulting distribution is to 1.0 and the larger the &lt;code&gt;shape2&lt;/code&gt; the closer it is to 0.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Some Note for Pareto Distribution</title>
      <link>https://chenxing.space/blog/some-note-for-pareto-distribution/</link>
      <pubDate>Sat, 16 Jul 2022 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/some-note-for-pareto-distribution/</guid>
      <description>&lt;h2 id=&#34;power-law-distribution&#34;&gt;Power Law Distribution&lt;/h2&gt;
&lt;div class=&#34;alert alert-note&#34;&gt;
  &lt;div&gt;
    &lt;p&gt;log-log-scale of cCDF showing you a &lt;strong&gt;straight line&lt;/strong&gt; ?&lt;/p&gt;
&lt;p&gt;This is the &lt;em&gt;signature of the Power Law distribution&lt;/em&gt;.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;h2 id=&#34;r-code&#34;&gt;R Code&lt;/h2&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;library&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;zetaEDA&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;library&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;zetaclv&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;enable_zeta_ggplot_theme&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;()&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# transactional data for cohort 2019&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;cohort19&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;eg_trans_data&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;with_groups&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;cust&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;mutate&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;    &lt;span class=&#34;n&#34;&gt;fp_yr&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;min&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;lubridate&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;::&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;year&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;date&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;filter&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;fp_yr&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;==&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;2019&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;select&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;o&#34;&gt;-&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;fp_yr&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;c1&#34;&gt;# build cbs data&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;dcbs&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;generate_cbs&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;cohort19&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;timeUnit&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;weeks&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;## Note that: time unit is in &amp;lt; weeks &amp;gt;
&lt;/code&gt;&lt;/pre&gt;&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;nf&#34;&gt;head&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;dcbs&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;##      cust x      t.x     litt sales sales.x      first     T.cal
## 1 uid0001 1 20.00000 2.995732  4644    1174 2019-12-02  79.00000
## 2 uid0005 0  0.00000 0.000000  1169       0 2019-08-08  95.57143
## 3 uid0006 1 50.71429 3.926208  1430     922 2019-04-20 111.28571
## 4 uid0010 0  0.00000 0.000000  2820       0 2019-02-15 120.42857
## 5 uid0011 0  0.00000 0.000000  6460       0 2019-01-15 124.85714
## 6 uid0012 0  0.00000 0.000000   473       0 2019-10-07  87.00000
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Note that &lt;code&gt;t.x&lt;/code&gt; is the &lt;em&gt;Time between first and last transactions&lt;/em&gt;. This is the &amp;ldquo;observed&amp;rdquo; part of lifetime. Let&amp;rsquo;s look at the distribution of &lt;code&gt;t.x&lt;/code&gt;.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; class=&#34;chroma&#34;&gt;&lt;code class=&#34;language-r&#34; data-lang=&#34;r&#34;&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;dtmp&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;lt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;dcbs&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;c1&#34;&gt;# remove single purchase customers&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;filter&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;t.x&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;&amp;gt;&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;0&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;c1&#34;&gt;# get value of cdf, P(X &amp;lt;= x)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;mutate&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;cdf&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;nf&#34;&gt;ecdf&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;t.x&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;t.x&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;c1&#34;&gt;# get ccef, P(X &amp;gt; x)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;mutate&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;ccdf&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;m&#34;&gt;1&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;-&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;cdf&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;&lt;span class=&#34;n&#34;&gt;dtmp&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;%&amp;gt;%&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;ggplot&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;nf&#34;&gt;aes&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;x&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;t.x&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;,&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;y&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;n&#34;&gt;ccdf&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;))&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;geom_point&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;n&#34;&gt;color&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;=&lt;/span&gt; &lt;span class=&#34;s&#34;&gt;&amp;#34;red&amp;#34;&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;)&lt;/span&gt; &lt;span class=&#34;o&#34;&gt;+&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;line&#34;&gt;&lt;span class=&#34;cl&#34;&gt;  &lt;span class=&#34;nf&#34;&gt;geom_line&lt;/span&gt;&lt;span class=&#34;p&#34;&gt;()&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;img src=&#34;https://chenxing.space/blog/some-note-for-pareto-distribution/index.en_files/figure-html/unnamed-chunk-1-1.png&#34; width=&#34;672&#34; /&gt;
&lt;h2 id=&#34;references&#34;&gt;References&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=9JkWtaVCqs0&amp;amp;t=751s&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Youtube video: Network Analysis. Lecture 2. Power laws.&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</description>
    </item>
    
    <item>
      <title>Theory Behind Pareto/NBD Part 1</title>
      <link>https://chenxing.space/blog/theory-behind-pnbd-prediciton-part-1/</link>
      <pubDate>Sun, 31 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/theory-behind-pnbd-prediciton-part-1/</guid>
      <description>

&lt;div id=&#34;TOC&#34;&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#introduction&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1&lt;/span&gt; Introduction&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#model-assumptions&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2&lt;/span&gt; PNBD Model Assumptions&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#two-stages-in-the-lifetime&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.1&lt;/span&gt; Two stages in the lifetime&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#poisson-purchase&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.2&lt;/span&gt; Poisson Purchase&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#exponential-lifetime&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.3&lt;/span&gt; Exponential Lifetime&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#gamma-transaction-rate&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.4&lt;/span&gt; Gamma transaction rate&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#gamma-dropout-rate&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.5&lt;/span&gt; Gamma dropout rate&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#two-processes-are-independent&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.6&lt;/span&gt; Two processes are Independent&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#why-named-paretonbd&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;3&lt;/span&gt; Why named Pareto/NBD?&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#poisson-gamma-mixture&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;3.1&lt;/span&gt; Poisson Gamma Mixture&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#exponential-gamma-mixture&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;3.2&lt;/span&gt; Exponential Gamma Mixture&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#reference&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;4&lt;/span&gt; Reference&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;

&lt;div id=&#34;introduction&#34; class=&#34;section level1&#34; number=&#34;1&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;1&lt;/span&gt; Introduction&lt;/h1&gt;
&lt;p&gt;The &lt;strong&gt;Pareto/NBD&lt;/strong&gt; model was developed by Schmittlein et al. (1987) to describe &lt;strong&gt;repeat-buying behavior&lt;/strong&gt; in a &lt;strong&gt;noncontractual&lt;/strong&gt; setting.&lt;/p&gt;
&lt;p&gt;There are 4 key questions:&lt;/p&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;How many “alive” customers does the firm now have?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How has this customer base grown over the past year?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Which individuals on this list most likely represent active customers? Inactive customers?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What level of transactions should be expected next year by those on the list, both individually and collectively?&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;In order to answer these questions, we need to build up the model(s) to estimate:&lt;/p&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;What is the &lt;span class=&#34;math inline&#34;&gt;\(\mathbb{P}(alive|\text{her trans infor})\)&lt;/span&gt;?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What is the &lt;span class=&#34;math inline&#34;&gt;\(\mathbb{E}(\text{# of trans}|\text{her trans infor})\)&lt;/span&gt;?&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
&lt;div id=&#34;model-assumptions&#34; class=&#34;section level1&#34; number=&#34;2&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;2&lt;/span&gt; PNBD Model Assumptions&lt;/h1&gt;
&lt;div id=&#34;two-stages-in-the-lifetime&#34; class=&#34;section level2&#34; number=&#34;2.1&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.1&lt;/span&gt; Two stages in the lifetime&lt;/h2&gt;
&lt;p&gt;Customers go through &lt;strong&gt;2 stages&lt;/strong&gt; in their “lifetime”: they are “&lt;strong&gt;alive&lt;/strong&gt;” for some period of time, then become permanently &lt;strong&gt;inactive&lt;/strong&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;poisson-purchase&#34; class=&#34;section level2&#34; number=&#34;2.2&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.2&lt;/span&gt; Poisson Purchase&lt;/h2&gt;
&lt;p&gt;Given a customer while alive, the number of transactions follows Poisson distribution with parameter &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt;, called &lt;strong&gt;transaction rate&lt;/strong&gt;. The probability of observing &lt;span class=&#34;math inline&#34;&gt;\(x\)&lt;/span&gt; transactions in the time interval &lt;span class=&#34;math inline&#34;&gt;\((0,t]\)&lt;/span&gt; is given by:&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[P(X(t) = x | \lambda ) =  e^{-\lambda t}\frac{(\lambda t)^x}{x!}, \ x = 0,1,2,...\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;This is equivalent to assuming that the time between transactions is &lt;span class=&#34;math inline&#34;&gt;\(Exp(\lambda)\)&lt;/span&gt;,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[f(t_j - t_{j-1} | \lambda ) = \lambda e^{(t_j - t_{j-1})}, \ t_j &amp;gt; t_{j-1}&amp;gt; 0,
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;where &lt;span class=&#34;math inline&#34;&gt;\(t_j\)&lt;/span&gt; is the time of the jth purchase.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;exponential-lifetime&#34; class=&#34;section level2&#34; number=&#34;2.3&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.3&lt;/span&gt; Exponential Lifetime&lt;/h2&gt;
&lt;p&gt;A customer’s unobserved “&lt;strong&gt;lifetime&lt;/strong&gt;” of length &lt;span class=&#34;math inline&#34;&gt;\(\tau\)&lt;/span&gt;, &lt;span class=&#34;math display&#34;&gt;\[\tau \sim Exp(\mu), \ \ f(\tau | \mu) = \mu e^{-\mu\tau},\]&lt;/span&gt; where &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt; is called &lt;strong&gt;dropout rate&lt;/strong&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;gamma-transaction-rate&#34; class=&#34;section level2&#34; number=&#34;2.4&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.4&lt;/span&gt; Gamma transaction rate&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity&lt;/strong&gt; in &lt;strong&gt;transaction rates&lt;/strong&gt; across customers follows a gamma distribution with shape parameter &lt;span class=&#34;math inline&#34;&gt;\(r\)&lt;/span&gt; and scale parameter &lt;span class=&#34;math inline&#34;&gt;\(\alpha\)&lt;/span&gt;:&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\lambda \sim Gamma(r, \alpha ), \ \ g(\lambda|r, \alpha ) = \frac{\alpha ^r \lambda^{r-1}e^{-\lambda \alpha }}{\Gamma (r)}.
\]&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;gamma-dropout-rate&#34; class=&#34;section level2&#34; number=&#34;2.5&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.5&lt;/span&gt; Gamma dropout rate&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Heterogeneity&lt;/strong&gt; in &lt;strong&gt;dropout rates&lt;/strong&gt; across customers follows a gamma distribution with shape parameter &lt;span class=&#34;math inline&#34;&gt;\(s\)&lt;/span&gt; and scale parameter &lt;span class=&#34;math inline&#34;&gt;\(\beta\)&lt;/span&gt;:&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\mu \sim Gamma(s, \beta), \ \ g(\mu | s, \beta ) = \frac{\beta^s\mu^{s-1}e^{-\mu\beta}}{\Gamma(s)}. \]&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;two-processes-are-independent&#34; class=&#34;section level2&#34; number=&#34;2.6&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.6&lt;/span&gt; Two processes are Independent&lt;/h2&gt;
&lt;p&gt;The transaction rate &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and the dropout rate &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt; vary &lt;strong&gt;independently&lt;/strong&gt; across customers,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\lambda \perp \mu.
\]&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;why-named-paretonbd&#34; class=&#34;section level1&#34; number=&#34;3&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;3&lt;/span&gt; Why named Pareto/NBD?&lt;/h1&gt;
&lt;p&gt;Short answer:&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\text{Poisson Purchase} + \text{Gamma transaction rate} \implies \text{NegBin}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\text{Exponential lifetime} + \text{Gamma dropout rate} \implies \text{Pareto}
\]&lt;/span&gt;&lt;/p&gt;
&lt;div id=&#34;poisson-gamma-mixture&#34; class=&#34;section level2&#34; number=&#34;3.1&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;3.1&lt;/span&gt; Poisson Gamma Mixture&lt;/h2&gt;
&lt;div class=&#34;theorem&#34;&gt;
&lt;p&gt;&lt;span id=&#34;thm:unnamed-chunk-2&#34; class=&#34;theorem&#34;&gt;&lt;strong&gt;Theorem 3.1  &lt;/strong&gt;&lt;/span&gt;If we assume the Poisson purchase and the Gamma transaction rate, then the distribution of the number of transactions while the customer is alive is Negative Binomial (NBD).&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&#34;proof&#34;&gt;
&lt;p&gt;&lt;span id=&#34;unlabeled-div-1&#34; class=&#34;proof&#34;&gt;&lt;em&gt;Proof&lt;/em&gt;. &lt;/span&gt;&lt;span class=&#34;math display&#34;&gt;\[\begin{align}
P(X(t) = x | r, \alpha ) &amp;amp;= \int_{0}^{\infty}P(X(t) = x | \lambda )g(\lambda |r, \alpha ) d\lambda \\
&amp;amp; = \int_{0}^{\infty} e^{-\lambda t}\frac{(\lambda t)^x}{x!}\frac{\lambda^{r-1}\alpha^re^{-\lambda \alpha }}{\Gamma(r)}  d\lambda\\

&amp;amp; = \frac{\alpha ^r}{\Gamma(r)}\frac{t^x}{x!}\int_{0}^{\infty }\lambda ^{x+r-1}e^{-\lambda (t+\alpha )}  d\lambda, \text{ let } u = (t+\alpha )\lambda ,\\

&amp;amp; = \frac{\alpha ^r}{\Gamma(r)}\frac{t^x}{x!}\frac{1}{(t+\alpha)^{x+r} }\int_{0}^{\infty }u^{x+r-1}e^{-u}du, \text{ note the form of } \Gamma(.),\\

&amp;amp; = \frac{\alpha ^r}{\Gamma(r)}\frac{t^x}{x!}\frac{1}{(t+\alpha)^{x+r} }\Gamma(x+r) \\

&amp;amp; = \frac{\Gamma(r+x)}{\Gamma(r)x!}\left ( \frac{\alpha }{\alpha +t} \right )^{r}\left ( \frac{t}{\alpha +t} \right )^x  

\end{align}\]&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Note that the last line is the density of &lt;a href=&#34;https://en.wikipedia.org/wiki/Negative_binomial_distribution%23Alternative_formulations&#34;&gt;negative binomial&lt;/a&gt;. It looks a little bit different from our familiar version of NegBin, and the parameter &lt;span class=&#34;math inline&#34;&gt;\(r\)&lt;/span&gt; extends to the &lt;span class=&#34;math inline&#34;&gt;\(\mathbb{R}^{+}\)&lt;/span&gt;. In this case, it is called &lt;strong&gt;Polya distribution&lt;/strong&gt; which is a special case of negative binomial.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;exponential-gamma-mixture&#34; class=&#34;section level2&#34; number=&#34;3.2&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;3.2&lt;/span&gt; Exponential Gamma Mixture&lt;/h2&gt;
&lt;div class=&#34;theorem&#34;&gt;
&lt;p&gt;&lt;span id=&#34;thm:unnamed-chunk-4&#34; class=&#34;theorem&#34;&gt;&lt;strong&gt;Theorem 3.2  &lt;/strong&gt;&lt;/span&gt;If we assume the Exponential lifetime and the Gamma dropout rate, then the distribution of the “lifetime” is “Pareto distribution of the second kind”.&lt;/p&gt;
&lt;/div&gt;
&lt;div class=&#34;proof&#34;&gt;
&lt;p&gt;&lt;span id=&#34;unlabeled-div-2&#34; class=&#34;proof&#34;&gt;&lt;em&gt;Proof&lt;/em&gt;. &lt;/span&gt;&lt;span class=&#34;math display&#34;&gt;\[\begin{align}
f(\tau|s, \beta ) &amp;amp;= \int_{0}^{\infty }f(\tau|\mu)g(\mu|s, \beta)d\mu \\
&amp;amp;=  \int_{0}^{\infty }\mu e^{-\mu\tau}\frac{\beta e^{-\beta\mu}(\beta\mu)^{s-1}}{\Gamma(s)}d\mu\\
&amp;amp;= \frac{\beta^s}{\Gamma(s)} \int_{0}^{\infty }\mu^{s} e^{-\mu(\tau+\beta)}d\mu, \text{ let } \ \  u = \mu(\tau+\beta) \\
&amp;amp;= \frac{\beta^s}{\Gamma(s)} \int_{0}^{\infty }\frac{1}{(\tau+\beta)^s}u^s e^{-u}\frac{1}{\tau+\beta}du\\
&amp;amp;= \frac{\beta^s}{\Gamma(s)}\frac{1}{(\tau+\beta)^{s+1}}\int_{0}^{\infty }u^s e^{-u}du\\
&amp;amp;= \frac{\beta^s}{\Gamma(s)}\frac{1}{(\tau+\beta)^{s+1}}\Gamma(s+1)\\
&amp;amp;= \frac{\beta^s}{\Gamma(s)}\frac{1}{(\tau+\beta)^{s+1}}s\Gamma(s)\\
&amp;amp;= \frac{s}{\beta}\left ( \frac{\beta}{\tau+\beta} \right )^{s+1}
\end{align}\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Note that&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\begin{align}
F(\tau|s, \beta ) &amp;amp;= \int_{0}^{\infty }F(\tau|\mu)g(\mu|s, \beta)d\mu \\
&amp;amp;= 1 - \left ( \frac{\beta}{\beta + \tau} \right ) ^{s}
\end{align}
\]&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Therefore, if we assume &lt;strong&gt;Exponential&lt;/strong&gt; &lt;strong&gt;lifetime&lt;/strong&gt; and &lt;strong&gt;Gamma&lt;/strong&gt; &lt;strong&gt;dropout rate&lt;/strong&gt;, we have ended with &lt;a href=&#34;https://en.wikipedia.org/wiki/Lomax_distribution&#34;&gt;Pareto Type II distribution&lt;/a&gt;, or more specifically, &lt;a href=&#34;https://en.wikipedia.org/wiki/Lomax_distribution&#34;&gt;Lomax distribution&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;In conclusion, the &lt;strong&gt;NBD&lt;/strong&gt; and &lt;strong&gt;Pareto&lt;/strong&gt; labels for each of the sub-models naturally leads to the name of the integrated model.&lt;/p&gt;
&lt;p&gt;In the next post we will talk about the likelihood, the mean of the Pareto/NBD model, and other related derivations, eg. probability of the customer being alive.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;reference&#34; class=&#34;section level1&#34; number=&#34;4&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;4&lt;/span&gt; Reference&lt;/h1&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;Schmittlein DC, Morrison DG, Colombo R (1987). “Counting Your Customers: Who-Are They and What Will They Do Next?” Management Science, 33(1), 1-24.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Fader PS, Hardie BGS (2005). “A Note on Deriving the Pareto/NBD Model and Related Expressions.” &lt;a href=&#34;http://www.brucehardie.com/notes/009/pareto_nbd_derivations_2005-11-05.pdf&#34;&gt;URL&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Fader PS, Hardie BGS (2007). “Incorporating time-invariant covariates into the Pareto/NBD and BG/NBD models.” &lt;a href=&#34;http://www.brucehardie.com/notes/019/time_invariant_covariates.pdf&#34;&gt;URL&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Fader PS, Hardie BGS (2020). “Deriving an Expression for P(X(t)=x) Under the Pareto/NBD Model.” &lt;a href=&#34;https://www.brucehardie.com/notes/012/pareto_NBD_pmf_derivation_rev.pdf&#34;&gt;URL&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Theory Behind Pareto/NBD Part 2</title>
      <link>https://chenxing.space/blog/theory-behond-pnbd-prediction-part-2/</link>
      <pubDate>Sun, 31 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/theory-behond-pnbd-prediction-part-2/</guid>
      <description>

&lt;div id=&#34;TOC&#34;&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#deriving-the-likelihood-function&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1&lt;/span&gt; Deriving the Likelihood Function&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#some-notation&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1.1&lt;/span&gt; Some notation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#conditional-on-lambda-and-mu&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1.2&lt;/span&gt; Conditional on &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#removing-the-conditioning-on-lambda-and-mu&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1.3&lt;/span&gt; Removing the Conditioning on &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#mle-for-r-alpha-s-beta&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1.4&lt;/span&gt; MLE for &lt;span class=&#34;math inline&#34;&gt;\(r, \alpha, s, \beta\)&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#derivations&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2&lt;/span&gt; Derivations&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#mean-of-the-paretonbd-model&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.1&lt;/span&gt; Mean of the Pareto/NBD Model&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#derivation-of-palive&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.2&lt;/span&gt; Derivation of PAlive&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#conditional-expectation-of-transactions&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.3&lt;/span&gt; Conditional Expectation of Transactions&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#reference&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;3&lt;/span&gt; Reference&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;

&lt;div id=&#34;deriving-the-likelihood-function&#34; class=&#34;section level1&#34; number=&#34;1&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;1&lt;/span&gt; Deriving the Likelihood Function&lt;/h1&gt;
&lt;p&gt;Last time we talked about the ParetoNBD Model. Today we’ll derive the model likelihood function.&lt;/p&gt;
&lt;div id=&#34;some-notation&#34; class=&#34;section level2&#34; number=&#34;1.1&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;1.1&lt;/span&gt; Some notation&lt;/h2&gt;
&lt;p&gt;For an customer,&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;images/transactionTime.png&#34; /&gt;&lt;/p&gt;
&lt;p&gt;Define:&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
x = \text{the number of purchase,}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
t_i = \text{the time of ith purchase}, \ \text{ where } 1 \le t_i \le t_x,
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
t_x = \text{the time of last purchase in the history,}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
T = \text{total time being observed}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Next, we’ll show that it is &lt;strong&gt;sufficient&lt;/strong&gt; to use individual’s &lt;span class=&#34;math inline&#34;&gt;\((x, t_x, T)\)&lt;/span&gt; to describe his/her purchase behavior in Pareto/NBD model.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;conditional-on-lambda-and-mu&#34; class=&#34;section level2&#34; number=&#34;1.2&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;1.2&lt;/span&gt; Conditional on &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;Assume a customer’s &lt;span class=&#34;math inline&#34;&gt;\(x\)&lt;/span&gt; transactions occurred during the period &lt;span class=&#34;math inline&#34;&gt;\((0,T]\)&lt;/span&gt;; we denote these times by &lt;span class=&#34;math inline&#34;&gt;\(t_1, t_2, t_3, \cdots, t_x\)&lt;/span&gt;.&lt;/p&gt;
&lt;p&gt;There are two possible ways this pattern of transactions could arise:&lt;/p&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;The customer is still alive at the end of the observation period (i.e., &lt;span class=&#34;math inline&#34;&gt;\(\tau &amp;gt; T\)&lt;/span&gt; ), the individual-level likelihood function is simply the product of the (inter-transaction-time) &lt;strong&gt;exponential&lt;/strong&gt; pdf and the associated survivor function:&lt;/li&gt;
&lt;/ol&gt;
&lt;span class=&#34;math display&#34;&gt;\[\begin{aligned}

L\left(\lambda \mid t_{1}, \ldots, t_{x}, T, \tau&amp;gt;T\right) &amp;amp;= \lambda e^{-\lambda t_{1}} \lambda e^{-\lambda\left(t_{2}-t_{1}\right)} \cdots \lambda e^{-\lambda\left(t_{x}-t_{x-1}\right)} e^{-\lambda\left(T-t_{x}\right)} \\

&amp;amp;=\lambda^{x} e^{-\lambda T}

\end{aligned}\]&lt;/span&gt;
&lt;ol start=&#34;2&#34; style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;The customer became &lt;strong&gt;inactive&lt;/strong&gt; at some time &lt;span class=&#34;math inline&#34;&gt;\(\tau\)&lt;/span&gt; in the interval &lt;span class=&#34;math inline&#34;&gt;\((t_x, T]\)&lt;/span&gt; (i.e. &lt;span class=&#34;math inline&#34;&gt;\(\tau \in (t_x, T]\)&lt;/span&gt;), in which case the individual-level likelihood function is&lt;/li&gt;
&lt;/ol&gt;
&lt;span class=&#34;math display&#34;&gt;\[\begin{aligned}

&amp;amp; L\left(\lambda \mid t_{1}, \ldots, t_{x}, T, \text { inactive at } \tau \in\left(t_{x}, T\right]\right) \\

&amp;amp;=\lambda e^{-\lambda t_{1}} \lambda e^{-\lambda\left(t_{2}-t_{1}\right)} \cdots \lambda e^{-\lambda\left(t_{x}-t_{x-1}\right)} e^{-\lambda\left(\tau-t_{x}\right)} \\

&amp;amp;=\lambda^{x} e^{-\lambda \tau}

\end{aligned}\]&lt;/span&gt;
&lt;p&gt;Note that in both cases, information on when each of the x transactions occurred is &lt;strong&gt;not required&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;We can replace &lt;span class=&#34;math inline&#34;&gt;\(t_1, ...t_x\)&lt;/span&gt; , &lt;span class=&#34;math inline&#34;&gt;\(T\)&lt;/span&gt; with &lt;span class=&#34;math inline&#34;&gt;\((x, t_x , T)\)&lt;/span&gt; where, by definition, &lt;span class=&#34;math inline&#34;&gt;\(t_x = 0\)&lt;/span&gt; when &lt;span class=&#34;math inline&#34;&gt;\(x = 0\)&lt;/span&gt;. In other words, &lt;span class=&#34;math inline&#34;&gt;\(t_x\)&lt;/span&gt;, &lt;span class=&#34;math inline&#34;&gt;\(x\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(T\)&lt;/span&gt; are &lt;strong&gt;sufficient&lt;/strong&gt; summaries of a customer’s transaction history. Using direct marketing terminology, &lt;span class=&#34;math inline&#34;&gt;\(t_x\)&lt;/span&gt; is &lt;strong&gt;recency&lt;/strong&gt; and &lt;span class=&#34;math inline&#34;&gt;\(x\)&lt;/span&gt; is &lt;strong&gt;frequency&lt;/strong&gt;.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;由以上两个事实可知，无需知晓客户每次的购买时间，&lt;strong&gt;第一次消费时间&lt;/strong&gt;、&lt;strong&gt;最后一次消费时间&lt;/strong&gt;、&lt;strong&gt;消费频次&lt;/strong&gt; 作为&lt;strong&gt;充分统计量&lt;/strong&gt;，已经足够我们导出似然函数了！&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Removing the conditioning on &lt;span class=&#34;math inline&#34;&gt;\(\tau\)&lt;/span&gt; gives us the following expression for the individual-level likelihood function:&lt;/p&gt;
&lt;span class=&#34;math display&#34;&gt;\[\begin{aligned}
L\left(\lambda, \mu \mid x, t_{x}, T\right)=&amp;amp; L(\lambda \mid x, T, \tau&amp;gt;T) P(\tau&amp;gt;T \mid \mu) + \\
&amp;amp;\int_{t_{x}}^{T} L\left(\lambda \mid x, T, \text { inactive at } \tau \in\left(t_{x}, T\right]\right) f(\tau \mid \mu) d \tau \\
&amp;amp;=\lambda^{x} e^{-\lambda T} e^{-\mu T}+\lambda^{x} \int_{t_{x}}^{T} e^{-\lambda \tau} \mu e^{-\mu \tau} d \tau \\
&amp;amp;=\lambda^{x} e^{-(\lambda+\mu) T}+\frac{\lambda^{x} \mu}{\lambda+\mu} e^{-(\lambda+\mu) t_{x}}-\frac{\lambda^{x} \mu}{\lambda+\mu} e^{-(\lambda+\mu) T} \\
&amp;amp;=\frac{\lambda^{x} \mu}{\lambda+\mu} e^{-(\lambda+\mu) t_{x}}+\frac{\lambda^{x+1}}{\lambda+\mu} e^{-(\lambda+\mu) T}
\end{aligned}\]&lt;/span&gt;
&lt;/div&gt;
&lt;div id=&#34;removing-the-conditioning-on-lambda-and-mu&#34; class=&#34;section level2&#34; number=&#34;1.3&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;1.3&lt;/span&gt; Removing the Conditioning on &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;We remove the conditioning on &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt; by taking the expectation of &lt;span class=&#34;math inline&#34;&gt;\(L(\lambda, \mu | x, t_x , T)\)&lt;/span&gt; over the distributions of &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt; :&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
L\left(r, \alpha, s, \beta \mid x, t_{x}, T\right)=\int_{0}^{\infty} \int_{0}^{\infty} L\left(\lambda, \mu \mid x, t_{x}, T\right) g(\lambda \mid r, \alpha) g(\mu \mid s, \beta) d \lambda d \mu
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;The computation is tedious, check the paper &lt;a href=&#34;http://www.brucehardie.com/notes/009/pareto_nbd_derivations_2005-11-05.pdf&#34;&gt;“A Note on Deriving the Pareto/NBD Model and Related Expressions”&lt;/a&gt; to know the details.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;mle-for-r-alpha-s-beta&#34; class=&#34;section level2&#34; number=&#34;1.4&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;1.4&lt;/span&gt; MLE for &lt;span class=&#34;math inline&#34;&gt;\(r, \alpha, s, \beta\)&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;Since we have derived the likelihood function &lt;span class=&#34;math inline&#34;&gt;\(L\left(r, \alpha, s, \beta \mid x, t_{x}, T\right)\)&lt;/span&gt;, the &lt;strong&gt;4&lt;/strong&gt; Pareto/NBD model parameters &lt;span class=&#34;math inline&#34;&gt;\((r, \alpha, s, \beta)\)&lt;/span&gt; can be estimated via the method of &lt;strong&gt;MLE&lt;/strong&gt;. Specifically, suppose we have a sample of &lt;span class=&#34;math inline&#34;&gt;\(N\)&lt;/span&gt; customers, where customer &lt;span class=&#34;math inline&#34;&gt;\(i\)&lt;/span&gt; had &lt;span class=&#34;math inline&#34;&gt;\(x_i\)&lt;/span&gt; transactions in the period &lt;span class=&#34;math inline&#34;&gt;\((0, T_i ]\)&lt;/span&gt;, with the last transaction occurring at &lt;span class=&#34;math inline&#34;&gt;\(t_{x_i}\)&lt;/span&gt; . The sample log-likelihood function is given by&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
L L(r, \alpha, s, \beta)=\sum_{i=1}^{N} \ln \left[L\left(r, \alpha, s, \beta \mid x_{i}, t_{x_{i}}, T_{i}\right)\right].
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;This can be maximized using standard numerical optimization routines. Therefore, we will obtain 4 &lt;strong&gt;maximum likelihood estimators&lt;/strong&gt; &lt;span class=&#34;math inline&#34;&gt;\((\widehat{r} \ , \ \widehat{\alpha} \ , \ \widehat{s} \ , \  \widehat{\beta})\)&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;derivations&#34; class=&#34;section level1&#34; number=&#34;2&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;2&lt;/span&gt; Derivations&lt;/h1&gt;
&lt;div id=&#34;mean-of-the-paretonbd-model&#34; class=&#34;section level2&#34; number=&#34;2.1&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.1&lt;/span&gt; Mean of the Pareto/NBD Model&lt;/h2&gt;
&lt;p&gt;Given that the number of transactions follows a Poisson process while the customer is alive,&lt;/p&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;if &lt;span class=&#34;math inline&#34;&gt;\(\tau &amp;gt; t\)&lt;/span&gt;, the expected number of transactions is simply &lt;span class=&#34;math inline&#34;&gt;\(\lambda t\)&lt;/span&gt;.&lt;/li&gt;
&lt;li&gt;if &lt;span class=&#34;math inline&#34;&gt;\(\tau \le t\)&lt;/span&gt;, the expected number of transactions in the interval (0, τ] is &lt;span class=&#34;math inline&#34;&gt;\(\lambda \tau\)&lt;/span&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Removing the conditioning on the time at which the customer becomes inactive, it follows that the expected number of transactions in the time interval &lt;span class=&#34;math inline&#34;&gt;\((0, t]\)&lt;/span&gt;, conditional on &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt;, is&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\begin{aligned}
E[X(t) \mid \lambda, \mu] &amp;amp;=\lambda t P(\tau&amp;gt;t \mid \mu)+\int_{0}^{t} \lambda \tau f(\tau \mid \mu) d \tau \\
&amp;amp;=\lambda t e^{-\mu t}+\lambda \int_{0}^{t} \mu \tau e^{-\mu \tau} d \tau \\
&amp;amp;=\lambda t e^{-\mu t}+\frac{\lambda}{\mu} \int_{0}^{t} \mu^{2} \tau e^{-\mu \tau} d \tau, \text{where integrand is an Erlang-2} \\
&amp;amp;=\lambda t e^{-\mu t}+\frac{\lambda}{\mu}\left\{1-e^{-\mu t}-\mu t e^{-\mu t}\right\} \\
&amp;amp;=\frac{\lambda}{\mu}-\frac{\lambda}{\mu} e^{-\mu t}
\end{aligned}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Now removing the Conditioning on &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt;,&lt;/p&gt;
&lt;span class=&#34;math display&#34; id=&#34;eq:popMean&#34;&gt;\[\begin{align}
E[X(t) \mid r, \alpha, s, \beta] &amp;amp;=\int_{0}^{\infty} \int_{0}^{\infty} E[X(t) \mid \lambda, \mu] g(\lambda \mid r, \alpha) g(\mu \mid s, \beta) d \lambda d \mu \\
&amp;amp;=\frac{r \beta}{\alpha(s-1)}-\frac{r \beta^{s}}{\alpha(s-1)(\beta+t)^{s-1}} \\
&amp;amp;=\frac{r \beta}{\alpha(s-1)}\left[1-\left(\frac{\beta}{\beta+t}\right)^{s-1}\right]
\tag{2.1}
\end{align}\]&lt;/span&gt;
&lt;/div&gt;
&lt;div id=&#34;derivation-of-palive&#34; class=&#34;section level2&#34; number=&#34;2.2&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.2&lt;/span&gt; Derivation of PAlive&lt;/h2&gt;
&lt;p&gt;The probability that a customer with purchase history &lt;span class=&#34;math inline&#34;&gt;\((x, t_x , T)\)&lt;/span&gt; is “alive” at time &lt;span class=&#34;math inline&#34;&gt;\(T\)&lt;/span&gt; is &lt;span class=&#34;math inline&#34;&gt;\(P(\tau &amp;gt; T)\)&lt;/span&gt;.&lt;/p&gt;
&lt;span class=&#34;math display&#34;&gt;\[\begin{aligned}
P\left(\tau&amp;gt;T \mid \lambda, \mu, x, t_{x}, T\right) &amp;amp;=\frac{L(\lambda \mid x, T, \tau&amp;gt;T) P(\tau&amp;gt;T \mid \mu)}{L\left(\lambda, \mu \mid x, t_{x}, T\right)} \\
&amp;amp;=\frac{\lambda^{x} e^{-(\lambda+\mu) T}}{L\left(\lambda, \mu \mid x, t_{x}, T\right)}
\end{aligned}\]&lt;/span&gt;
&lt;p&gt;As the &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt; are unobserved, we compute &lt;span class=&#34;math inline&#34;&gt;\(P(alive | x, t_x , T)\)&lt;/span&gt; for a randomly-chosen individual by taking the expectation of the above result over the distribution of &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt; , updated to take account of the information &lt;span class=&#34;math inline&#34;&gt;\((x, t_x , T)\)&lt;/span&gt;:&lt;/p&gt;
&lt;span class=&#34;math display&#34;&gt;\[\begin{array}{l}
P\left(\text { alive } \mid r, \alpha, s, \beta, x, t_{x}, T\right) \\
\qquad=\int_{0}^{\infty} \int_{0}^{\infty} P\left(\tau&amp;gt;T \mid \lambda, \mu, x, t_{x}, T\right) g\left(\lambda, \mu \mid r, \alpha, s, \beta, x, t_{x}, T\right) d \lambda d \mu
\end{array}\]&lt;/span&gt;
&lt;p&gt;By Bayes’ theorem, the joint posterior distribution of &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt; is&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
g\left(\lambda, \mu \mid r, \alpha, s, \beta, x, t_{x}, T\right)=\frac{L\left(\lambda, \mu \mid x, t_{x}, T\right) g(\lambda \mid r, \alpha) g(\mu \mid s, \beta)}{L\left(r, \alpha, s, \beta \mid x, t_{x}, T\right)}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Thus,&lt;/p&gt;
&lt;span class=&#34;math display&#34;&gt;\[\begin{array}{l}
P\left(\text { alive } \mid r, \alpha, s, \beta, x, t_{x}, T\right) \\
\quad=\int_{0}^{\infty} \int_{0}^{\infty} \lambda^{x} e^{-(\lambda+\mu) T} g(\lambda \mid r, \alpha) g(\mu \mid s, \beta) d \lambda d \mu / L\left(r, \alpha, s, \beta \mid x, t_{x}, T\right) \\
\quad=\frac{\Gamma(r+x) \alpha^{r} \beta^{s}}{\Gamma(r)(\alpha+T)^{r+x}(\beta+T)^{s}} / L\left(r, \alpha, s, \beta \mid x, t_{x}, T\right)\\
\quad=\left\{1+\left(\frac{s}{r+s+x}\right)(\alpha+T)^{r+x}(\beta+T)^{s} \mathrm{~A}_{0}\right\}^{-1}
\end{array}\]&lt;/span&gt;
&lt;p&gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202207091841479.png&#34; alt=&#34;cap2022-07-09 18.38.17&#34; style=&#34;zoom:30%;&#34; /&gt;&lt;/p&gt;
&lt;p&gt;For details check the reference paper. Note that, the above result is the formula to calculate &lt;strong&gt;PAlive&lt;/strong&gt; used in &lt;code&gt;BTYD&lt;/code&gt; 📦 implemented in R.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;conditional-expectation-of-transactions&#34; class=&#34;section level2&#34; number=&#34;2.3&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.3&lt;/span&gt; Conditional Expectation of Transactions&lt;/h2&gt;
&lt;p&gt;Let random variable &lt;span class=&#34;math inline&#34;&gt;\(Y(t) = \text{num of purchase made in } (T, T+t]\)&lt;/span&gt;. We are interested in computing &lt;span class=&#34;math inline&#34;&gt;\(E(Y(t)|x, t_x, T)\)&lt;/span&gt;, the expected number of purchase in the period &lt;span class=&#34;math inline&#34;&gt;\((T, T+t]\)&lt;/span&gt; for a customer with purchase history &lt;span class=&#34;math inline&#34;&gt;\((x, t_x, T)\)&lt;/span&gt;.&lt;/p&gt;
&lt;p&gt;If the customer is active at &lt;span class=&#34;math inline&#34;&gt;\(T\)&lt;/span&gt;,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\begin{array}{l}
&amp;amp;E[Y(t) \mid \lambda, \mu, \text { alive at } T]\\
&amp;amp;=\lambda t P(\tau&amp;gt;T+t \mid \mu, \tau&amp;gt;T)+\int_{T}^{T+t} \lambda \tau f(\tau \mid \mu, \tau&amp;gt;T) d \tau\\
&amp;amp;=\lambda t e^{-\mu t}+\lambda \int_{0}^{t} \mu \tau e^{-\mu \tau} d \tau \\
&amp;amp;=\frac{\lambda}{\mu}-\frac{\lambda}{\mu} e^{-\mu t}
\end{array}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Of course we don’t know whether a customer is alive at &lt;span class=&#34;math inline&#34;&gt;\(T\)&lt;/span&gt;; therefore&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
E\left[Y(t) \mid \lambda, \mu, x, t_{x}, T\right]=E[Y(t) \mid \lambda, \mu, \text { alive at } T] P\left(\tau&amp;gt;T \mid \lambda, \mu, x, t_{x}, T\right)
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Also, since &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(\mu\)&lt;/span&gt; are unobserved, we need to integrate them out:&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\begin{array}{c}
E\left[Y(t) \mid r, \alpha, s, \beta, x, t_{x}, T\right]=\int_{0}^{\infty} \int_{0}^{\infty}\left\{E[Y(t) \mid \lambda, \mu, \text { alive at } T] P\left(\tau&amp;gt;T \mid \lambda, \mu, x, t_{x}, T\right)\right. \\
\left.g\left(\lambda, \mu \mid r, \alpha, s, \beta, x, t_{x}, T\right)\right\} d \lambda d \mu
\end{array}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;After the tedious computation, we will get&lt;/p&gt;
&lt;span class=&#34;math display&#34;&gt;\[\begin{aligned}
&amp;amp;E\left[Y(t) \mid r, \alpha, s, \beta, x, t_{x}, T\right]\\

&amp;amp;=\{\frac{\Gamma(r+x) \alpha^{r} \beta^{s}}{\Gamma(r)(\alpha+T)^{r+x}(\beta+T)^{s}} / L\left(r, \alpha, s, \beta \mid x, t_{x}, T\right)\} \\
&amp;amp;\times \frac{(r+x)(\beta+T)}{(\alpha+T)(s-1)}\left[1-\left(\frac{\beta+T}{\beta+T+t}\right)^{s-1}\right]\\
&amp;amp;= \{P(\text{alive}|x, t_x, T)\} \times \text{updated mean of Pareto/NBD}
\end{aligned}\]&lt;/span&gt;
&lt;p&gt;Note that:&lt;/p&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;The first part, the bracketed term, is out expression for &lt;span class=&#34;math inline&#34;&gt;\(P(\text{alive}|x, t_x, T)\)&lt;/span&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The rest part is mean of the Pareto/NBD &lt;a href=&#34;#eq:popMean&#34;&gt;(2.1)&lt;/a&gt;, with “&lt;strong&gt;updated&lt;/strong&gt;” parameters that reflect the &lt;em&gt;individual’s behavior&lt;/em&gt; up to time &lt;span class=&#34;math inline&#34;&gt;\(T\)&lt;/span&gt; (assuming no “death” in &lt;span class=&#34;math inline&#34;&gt;\((0,T])\)&lt;/span&gt;).&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Next time, we’ll finally take about the prediction of CLV.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;reference&#34; class=&#34;section level1&#34; number=&#34;3&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;3&lt;/span&gt; Reference&lt;/h1&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;Schmittlein DC, Morrison DG, Colombo R (1987). “Counting Your Customers: Who-Are They and What Will They Do Next?” Management Science, 33(1), 1-24.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Fader PS, Hardie BGS (2005). “A Note on Deriving the Pareto/NBD Model and Related Expressions.” &lt;a href=&#34;http://www.brucehardie.com/notes/009/pareto_nbd_derivations_2005-11-05.pdf&#34;&gt;URL&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Fader PS, Hardie BGS (2007). “Incorporating time-invariant covariates into the Pareto/NBD and BG/NBD models.” &lt;a href=&#34;http://www.brucehardie.com/notes/019/time_invariant_covariates.pdf&#34;&gt;URL&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Fader PS, Hardie BGS (2020). “Deriving an Expression for P(X(t)=x) Under the Pareto/NBD Model.” &lt;a href=&#34;https://www.brucehardie.com/notes/012/pareto_NBD_pmf_derivation_rev.pdf&#34;&gt;URL&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Gamma-Gamma Spend Model</title>
      <link>https://chenxing.space/blog/note-for-gamma-gamma-model/</link>
      <pubDate>Sat, 30 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/note-for-gamma-gamma-model/</guid>
      <description>

&lt;div id=&#34;TOC&#34;&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#model-assumption&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1&lt;/span&gt; Model Assumption&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#compute-ez-mid-barz-x&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2&lt;/span&gt; Compute &lt;span class=&#34;math inline&#34;&gt;\(E(Z \mid \bar{z}, x)\)&lt;/span&gt;&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#derive-related-conditional-density&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.1&lt;/span&gt; Derive related conditional density&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#get-the-result&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.2&lt;/span&gt; Get the result&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#understand-the-result&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;3&lt;/span&gt; Understand the result&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#key-point&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;3.1&lt;/span&gt; Key point&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#references&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;4&lt;/span&gt; References&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;How to predict a customer’s &lt;strong&gt;mean spending&lt;/strong&gt; in the future?&lt;/p&gt;
&lt;p&gt;Answer: You can use gamma-gamma model.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;div id=&#34;model-assumption&#34; class=&#34;section level1&#34; number=&#34;1&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;1&lt;/span&gt; Model Assumption&lt;/h1&gt;
&lt;p&gt;There are &lt;strong&gt;3&lt;/strong&gt; general assumptions for this model:&lt;/p&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;The monetary value (e.g. $, ¥) of a customer’s given transaction &lt;strong&gt;varies randomly around their average&lt;/strong&gt; transaction value.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Average transaction values vary across customers but &lt;strong&gt;do not vary over time&lt;/strong&gt; for any given individual.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The distribution of average transaction values across customers is &lt;strong&gt;independent of the transaction process&lt;/strong&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For a customer with &lt;span class=&#34;math inline&#34;&gt;\(x\)&lt;/span&gt; transactions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;p&gt;let &lt;span class=&#34;math inline&#34;&gt;\(\{z_1, z_2, \cdots, z_x\}\)&lt;/span&gt; denote the value of each transaction.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;span class=&#34;math inline&#34;&gt;\(z_i\)&lt;/span&gt;’s are samples from the distribution of R.V. &lt;span class=&#34;math inline&#34;&gt;\(Z\)&lt;/span&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The customer’s observed average transaction value is &lt;span class=&#34;math inline&#34;&gt;\(\bar{z} = \sum_{i = 1}^{x}\frac{z_i}{x}\)&lt;/span&gt;.&lt;/p&gt;
&lt;p&gt;However, &lt;span class=&#34;math inline&#34;&gt;\(\bar{z}\)&lt;/span&gt; is an &lt;strong&gt;imperfect estimate&lt;/strong&gt; of the &lt;strong&gt;unobserved&lt;/strong&gt; mean transaction value.&lt;/p&gt;
&lt;p&gt;Why? Consider when a customer only had very limited transactions, say &lt;em&gt;1 or 2&lt;/em&gt; purchases, then it is questionable to use his average spending to estimate the spending power. At least we should use the population mean as the standard criteria to help. On the other hand, if the customer had enough purchase history, then we want to emphasize more on his own average spending while put relatively less weight on the population mean.&lt;/p&gt;
&lt;p&gt;The goal is to make inference about&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
E(Z| \bar{z}, x)
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;What is distribution for &lt;span class=&#34;math inline&#34;&gt;\(Z\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;Maybe log-normal or gamma, since spend data tend to be &lt;strong&gt;right skewed&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Here we assume the &lt;strong&gt;gamma distribution&lt;/strong&gt;. More formally, we assume that:&lt;/p&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;&lt;span class=&#34;math inline&#34;&gt;\(Z \sim gamma(p, \nu)\)&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;span class=&#34;math inline&#34;&gt;\(\nu \sim gamma(q, \gamma)\)&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;From above setting, we know&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
E(Z|p, \nu) := \zeta = \frac{p}{\nu}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\text{total spend} = \sum_{i = 1}^{x}z_i \sim gamma(px, \nu)
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\text{average spend} = \bar{z} = \sum_{i = 1}^{x}\frac{z_i}{x} \sim gamma(px, \nu x)
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;They can be easily proved using MGF.&lt;/p&gt;
&lt;p&gt;This results in what we call the &lt;strong&gt;gamma-gamma&lt;/strong&gt; model of spend.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;compute-ez-mid-barz-x&#34; class=&#34;section level1&#34; number=&#34;2&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;2&lt;/span&gt; Compute &lt;span class=&#34;math inline&#34;&gt;\(E(Z \mid \bar{z}, x)\)&lt;/span&gt;&lt;/h1&gt;
&lt;p&gt;We wish to make inferences about an individual customer’s mean spending given &lt;span class=&#34;math inline&#34;&gt;\(\bar{z}\)&lt;/span&gt;, which we denote as &lt;span class=&#34;math inline&#34;&gt;\(E(Z|\bar{z}, x)\)&lt;/span&gt;.&lt;/p&gt;
&lt;p&gt;Note that,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34; id=&#34;eq:res&#34;&gt;\[
\begin{equation}
E(Z \mid \bar{z}, x) = E(Z \mid p, q, \gamma ; \bar{z}, x) =\int_{0}^{\infty }E(Z\mid p, \nu)g(\nu \mid p, q, \gamma, \bar{z},x )d\nu
\tag{2.1}
\end{equation}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
Z \sim gamma(p, v) \implies E(Z \mid p, \nu) = \frac{p}{\nu}
\]&lt;/span&gt;&lt;/p&gt;
&lt;div id=&#34;derive-related-conditional-density&#34; class=&#34;section level2&#34; number=&#34;2.1&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.1&lt;/span&gt; Derive related conditional density&lt;/h2&gt;
&lt;p&gt;By Bayes Theorem,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34; id=&#34;eq:1&#34;&gt;\[
\begin{equation}
g(\nu \mid p, q, \gamma ; \bar{z}, x) =\frac{f(\bar{z} \mid p, \nu ; x) g(\nu \mid q, \gamma)}{f(\bar{z} \mid p, q, \gamma ; x)}
\tag{2.2}
\end{equation}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;By assumption,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34; id=&#34;eq:2&#34;&gt;\[
\begin{equation}
\nu \sim gamma(q, \gamma) \ \ , \ g(\nu \mid q, \gamma) = \frac{\gamma^{q} \nu^{q-1} e^{-\gamma \nu}}{\Gamma(q)}
\tag{2.3}
\end{equation}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34; id=&#34;eq:3&#34;&gt;\[
\begin{equation}
\bar{z} \sim gamma(px, \nu x) \ \ , \ f(\bar{z} \mid p, \nu ; x) = \frac{(\nu x)^{p x} \bar{z}^{p x-1} e^{-\nu x \bar{z}}}{\Gamma(p x)}
\tag{2.4}
\end{equation}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Only need to derive &lt;span class=&#34;math inline&#34;&gt;\(f(\bar{z} \mid p, q, \gamma ; x)\)&lt;/span&gt;,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34; id=&#34;eq:4&#34;&gt;\[
\begin{align}
f(\bar{z} \mid p, q, \gamma ; x) &amp;amp;= \int_{0}^{\infty} f(\bar{z} \mid p, \nu; x) g(\nu \mid q, \gamma ) d\nu\\
&amp;amp;=\int_{0}^{\infty} \frac{(\nu x)^{p x} \bar{z}^{p x-1} e^{-\nu x \bar{z}}}{\Gamma(p x)} \frac{\gamma^{q} \nu^{q-1} e^{-\gamma \nu}}{\Gamma(q)} d \nu \\
&amp;amp;=\frac{\bar{z}^{p x-1} x^{p x} \gamma^{q}}{\Gamma(p x) \Gamma(q)} \int_{0}^{\infty} \nu^{p x+q-1} e^{-(\gamma+x \bar{z}) \nu} d \nu \\
&amp;amp;=\frac{\Gamma(p x+q)}{\Gamma(p x) \Gamma(q)} \frac{\bar{z}^{p x-1} x^{p x} \gamma^{q}}{(\gamma+x \bar{z})^{p x+q}} \\
&amp;amp;=\frac{1}{\bar{z} B(p x, q)}\left(\frac{\gamma}{\gamma+x \bar{z}}\right)^{q}\left(\frac{x \bar{z}}{\gamma+x \bar{z}}\right)^{p x}
\tag{2.5}
\end{align}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;We already have all we need to derive &lt;a href=&#34;#eq:1&#34;&gt;(2.2)&lt;/a&gt;, thus&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\begin{aligned}
g(\nu \mid p, q, \gamma ; \bar{z}, x) &amp;amp;=\frac{f(\bar{z} \mid p, \nu ; x) g(\nu \mid q, \gamma)}{f(\bar{z} \mid p, q, \gamma ; x)} \\
&amp;amp;=\frac{(\nu x)^{p x} \bar{z}^{p x-1} e^{-\nu x \bar{z}}}{\Gamma(p x)} \frac{\gamma^{q} \nu^{q-1} e^{-\gamma \nu}}{\Gamma(q)} / \frac{\Gamma(p x+q)}{\Gamma(p x) \Gamma(q)} \frac{\bar{z}^{p x-1} x^{p x} \gamma^{q}}{(\gamma+x \bar{z})^{p x+q}} \\
&amp;amp;=\frac{(\gamma+x \bar{z})^{p x+q} \nu^{p x+q-1} e^{-\nu(\gamma+x \bar{z})}}{\Gamma(p x+q)} \\
&amp;amp;\sim gamma(px+q, \gamma + x \bar{z})
\end{aligned}
\]&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;get-the-result&#34; class=&#34;section level2&#34; number=&#34;2.2&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.2&lt;/span&gt; Get the result&lt;/h2&gt;
&lt;p&gt;Now we have everything we need to derive &lt;a href=&#34;#eq:res&#34;&gt;(2.1)&lt;/a&gt;. Before plug in, let’s review the &lt;a href=&#34;https://en.wikipedia.org/wiki/Inverse-gamma_distribution&#34;&gt;inverse gamma distribution&lt;/a&gt; which will be used later.&lt;/p&gt;
&lt;p&gt;If &lt;span class=&#34;math inline&#34;&gt;\(X \sim Gamma(\alpha, \beta)\)&lt;/span&gt; then &lt;span class=&#34;math inline&#34;&gt;\(Y :=\frac{1}{x} \sim \text{Inv-Gamma}(\alpha, \beta)\)&lt;/span&gt;, with&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math inline&#34;&gt;\(E(Y) = \frac {\beta }{\alpha -1}\)&lt;/span&gt;, for &lt;span class=&#34;math inline&#34;&gt;\(\alpha &amp;gt; 1\)&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math inline&#34;&gt;\(Var(Y) = \frac {\beta ^{2}}{(\alpha -1)^{2}(\alpha -2)}\)&lt;/span&gt;, for &lt;span class=&#34;math inline&#34;&gt;\(\alpha &amp;gt; 2\)&lt;/span&gt;.&lt;/p&gt;
&lt;p&gt;Plug into &lt;a href=&#34;#eq:res&#34;&gt;(2.1)&lt;/a&gt;,&lt;/p&gt;
&lt;span class=&#34;math display&#34; id=&#34;eq:final-concise-form&#34;&gt;\[\begin{align}
E(Z \mid p, q, \gamma ; \bar{z}, x) &amp;amp;=\int_{0}^{\infty }E(Z|p, \nu)g(\nu |p, q, \gamma, \bar{z},x )d\nu \\
&amp;amp;= \int_{0}^{\infty } \frac{p}{\nu} g(\nu |p, q, \gamma, \bar{z},x )d\nu \\
&amp;amp;= p \cdot E(\frac{1}{\nu}|p, q, \gamma, \bar{z},x ), \\
&amp;amp;\text{note: }\\
&amp;amp;\frac{1}{\nu} \sim \text{Inv-Gamma}(px+q, \gamma+x \bar{z}) \\
&amp;amp;E(\text{Inv-Gamma}) = \frac{\text{scale}}{\text{shape - 1}} \\
&amp;amp;=\frac{p(\gamma+x \bar{z})}{p x+q-1}
\tag{2.6}
\end{align}\]&lt;/span&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;understand-the-result&#34; class=&#34;section level1&#34; number=&#34;3&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;3&lt;/span&gt; Understand the result&lt;/h1&gt;
&lt;p&gt;How to understand &lt;span class=&#34;math inline&#34;&gt;\(E(Z \mid p, q, \gamma ; \bar{z}, x)\)&lt;/span&gt;? What can we say about the estimator of the individual’s future spend per transaction given his/her historical average spend &lt;span class=&#34;math inline&#34;&gt;\(\bar{z}\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;First, let’s derive &lt;span class=&#34;math inline&#34;&gt;\(E(Z \mid p, q, \gamma)\)&lt;/span&gt;, which is the &lt;strong&gt;population mean&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\begin{aligned}
E(Z \mid p, q, \gamma) &amp;amp;= \int_{0}^{\infty }E(Z\mid p, \nu)g(\nu \mid q, \gamma)d\nu \\
&amp;amp;= \int_{0}^{\infty } \frac{p}{\nu}g(\nu \mid q, \gamma)d\nu\\
&amp;amp;= pE(\frac{1}{\nu}), \text{ where } \frac{1}{\nu} \sim \text{Inv-Gamma}(q, \gamma)\\
&amp;amp;= \frac{p\gamma}{q-1}
\end{aligned}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Now let’s rearrange the result of &lt;a href=&#34;#eq:final-concise-form&#34;&gt;(2.6)&lt;/a&gt;,&lt;/p&gt;
&lt;span class=&#34;math display&#34;&gt;\[\begin{align}
E(Z \mid p, q, \gamma ; \bar{z}, x) &amp;amp;= \frac{p(\gamma+x \bar{z})}{px+q-1} \\
&amp;amp;=\left(\frac{q-1}{p x+q-1}\right) \underbrace{\frac{p \gamma}{q-1}}_{\text{population mean}} + \left(\frac{p x}{p x+q-1}\right) \underbrace{\bar{z} \vphantom{\frac{p \gamma}{q-1}}}_{\text{observed average}}
\end{align}\]&lt;/span&gt;
&lt;div id=&#34;key-point&#34; class=&#34;section level2&#34; number=&#34;3.1&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;3.1&lt;/span&gt; Key point&lt;/h2&gt;
&lt;p&gt;We note that this is the &lt;strong&gt;weighted average&lt;/strong&gt; of the &lt;strong&gt;population mean&lt;/strong&gt;, &lt;span class=&#34;math inline&#34;&gt;\(E(Z|p, q, \gamma)\)&lt;/span&gt;, and the given individual’s &lt;strong&gt;observed average&lt;/strong&gt; transaction value, &lt;span class=&#34;math inline&#34;&gt;\(\bar{z}\)&lt;/span&gt;.&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
E(Z \mid \bar{z}, x) = w \cdot \{\text{population mean}\} + (1-w) \cdot \{\text{observed average}\}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;As the number of observations &lt;span class=&#34;math inline&#34;&gt;\((x)\)&lt;/span&gt; used to compute &lt;span class=&#34;math inline&#34;&gt;\(\bar{z}\)&lt;/span&gt; &lt;strong&gt;increases&lt;/strong&gt;, &lt;strong&gt;less&lt;/strong&gt; weight is placed on the &lt;strong&gt;population mean&lt;/strong&gt; and &lt;strong&gt;more&lt;/strong&gt; weight is placed on the customer’s &lt;strong&gt;observed average&lt;/strong&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;references&#34; class=&#34;section level1&#34; number=&#34;4&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;4&lt;/span&gt; References&lt;/h1&gt;
&lt;p&gt;[1] Fader PS, Hardie BG (2013). “The Gamma-Gamma Model of Monetary Value.” URL &lt;a href=&#34;http://www.brucehardie.com/notes/025/gamma_gamma.pdf&#34; class=&#34;uri&#34;&gt;http://www.brucehardie.com/notes/025/gamma_gamma.pdf&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;[2] Colombo R, Jiang W (1999). “A stochastic RFM model.” Journal of Interactive Marketing, 13(3), 2-12.&lt;/p&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Note for Beta Distribution</title>
      <link>https://chenxing.space/blog/note-for-beta-distribution/</link>
      <pubDate>Sat, 30 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/note-for-beta-distribution/</guid>
      <description>

&lt;div id=&#34;TOC&#34;&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#why-beta-distribution&#34; id=&#34;toc-why-beta-distribution&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1&lt;/span&gt; Why Beta Distribution?&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#model-probabilities&#34; id=&#34;toc-model-probabilities&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1.1&lt;/span&gt; Model probabilities&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#generalization-of-uniform&#34; id=&#34;toc-generalization-of-uniform&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;1.2&lt;/span&gt; Generalization of uniform&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#construction&#34; id=&#34;toc-construction&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2&lt;/span&gt; Construction&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#bank-and-post-office-story&#34; id=&#34;toc-bank-and-post-office-story&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.1&lt;/span&gt; Bank and Post Office Story&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#summary&#34; id=&#34;toc-summary&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.1.1&lt;/span&gt; Summary&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#plots&#34; id=&#34;toc-plots&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;2.2&lt;/span&gt; plots&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#reference&#34; id=&#34;toc-reference&#34;&gt;&lt;span class=&#34;toc-section-number&#34;&gt;3&lt;/span&gt; Reference&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;

&lt;p&gt;This post is out of date, please check the new post named “&lt;strong&gt;Beta Distribution — Intuition, Derivation, and Examples&lt;/strong&gt;”.&lt;/p&gt;
&lt;div id=&#34;why-beta-distribution&#34; class=&#34;section level1&#34; number=&#34;1&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;1&lt;/span&gt; Why Beta Distribution?&lt;/h1&gt;
&lt;div id=&#34;model-probabilities&#34; class=&#34;section level2&#34; number=&#34;1.1&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;1.1&lt;/span&gt; Model probabilities&lt;/h2&gt;
&lt;p&gt;The short story is that the Beta distribution can be understood as representing a distribution &lt;em&gt;of probabilities&lt;/em&gt;, that is, it represents all the possible values of a probability when we don’t know what that probability is.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;generalization-of-uniform&#34; class=&#34;section level2&#34; number=&#34;1.2&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;1.2&lt;/span&gt; Generalization of uniform&lt;/h2&gt;
&lt;p&gt;Give me a &lt;strong&gt;continuous&lt;/strong&gt; and &lt;strong&gt;bounded&lt;/strong&gt; random variable, em, except the &lt;em&gt;Uniform distribution&lt;/em&gt;. That is another way to look at &lt;em&gt;beta distribution&lt;/em&gt;, continuous and bounded between 0, 1; also the density is not flat.
&lt;span class=&#34;math display&#34;&gt;\[
X \sim Beta(a, b), \text{ where } a&amp;gt;0, \ b&amp;gt;0.
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
f_X(x) = c \cdot x ^{a-1}(1-x)^{b-1}, \text{ where } x&amp;gt;0.
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;What is &lt;span class=&#34;math inline&#34;&gt;\(c\)&lt;/span&gt;? Just a normalization constant. We’ll get the value of &lt;span class=&#34;math inline&#34;&gt;\(c\)&lt;/span&gt; later.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;construction&#34; class=&#34;section level1&#34; number=&#34;2&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;2&lt;/span&gt; Construction&lt;/h1&gt;
&lt;div id=&#34;bank-and-post-office-story&#34; class=&#34;section level2&#34; number=&#34;2.1&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.1&lt;/span&gt; Bank and Post Office Story&lt;/h2&gt;
&lt;p&gt;Let &lt;span class=&#34;math inline&#34;&gt;\(X\)&lt;/span&gt; be the waiting time at Bank,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
X \sim Gamma(n_1, \lambda)
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Let &lt;span class=&#34;math inline&#34;&gt;\(Y\)&lt;/span&gt; be the waiting time at Post Office,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
Y \sim Gamma(n_2, \lambda)
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Assume &lt;span class=&#34;math inline&#34;&gt;\(X\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(Y\)&lt;/span&gt; are independent.&lt;/p&gt;
&lt;p&gt;Then, what is the distribution of the proportion &lt;span class=&#34;math inline&#34;&gt;\(\frac{X}{X+Y}\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;Define &lt;span class=&#34;math inline&#34;&gt;\(T = X+Y\)&lt;/span&gt; be the total waiting time.&lt;/p&gt;
&lt;p&gt;Clearly, &lt;span class=&#34;math inline&#34;&gt;\(T \sim Gamma(n_1+n_2, \lambda)\)&lt;/span&gt;, proved by MGF.&lt;/p&gt;
&lt;p&gt;Define &lt;span class=&#34;math inline&#34;&gt;\(W = \frac{X}{X+Y}\)&lt;/span&gt; , the proportion of waiting time at Bank to the total waiting time.&lt;/p&gt;
&lt;p&gt;What is the distribution of &lt;span class=&#34;math inline&#34;&gt;\(W\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;We need to derive &lt;span class=&#34;math inline&#34;&gt;\(f_W(w)\)&lt;/span&gt;, but first let’s find the joint pdf &lt;span class=&#34;math inline&#34;&gt;\(f_{T,W}(t,w)\)&lt;/span&gt;
&lt;span class=&#34;math display&#34;&gt;\[
\begin{align}
    f_{T,W}(t,w) &amp;amp;= f_{X,Y}(x,y) \left | \frac{\partial(x,y)}{\partial(t,w)} \right|\\
    &amp;amp;= \frac{1}{\Gamma(n_1)}\lambda^{n_1}x^{n_1 - 1}e^{-\lambda x} \frac{1}{\Gamma(n_2)}\lambda^{n_2}x^{n_2 - 1}e^{-\lambda y} \left|-t\right| \\
    &amp;amp;= \lambda^{n_1+n_2}t^{n_1+n_2-1}e^{\lambda t} \frac{1}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} \\
    &amp;amp;= \frac{\lambda^{n_1+n_2}t^{n_1+n_2-1}e^{\lambda t}}{\Gamma(n_1+n_2)} \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}\\
    &amp;amp;= f_T(t) \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}
\end{align}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Then we find the marginal,&lt;/p&gt;
&lt;span class=&#34;math display&#34;&gt;\[\begin{aligned}
f_W(w) &amp;amp;= \int_0^\infty f_{T,W}(t,w) dt \\

&amp;amp;= \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} \cdot\int_0^\infty f_T(t)dt \\

&amp;amp;= \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1}
\end{aligned}\]&lt;/span&gt;
&lt;p&gt;Since &lt;span class=&#34;math inline&#34;&gt;\(f_W(w)\)&lt;/span&gt; is the pdf needed to be integrated to 1,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\int_0^1\frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)}w^{n_1 - 1}(1-w)^{n_2-1} dw \equiv 1
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;so the normalization constant should be&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
c = \frac{\Gamma(n_1+n_2)}{\Gamma(n_1)\Gamma(n_2)} := \frac{1}{B(n_1, n_2)}
\]&lt;/span&gt;&lt;/p&gt;
&lt;div id=&#34;summary&#34; class=&#34;section level3&#34; number=&#34;2.1.1&#34;&gt;
&lt;h3&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.1.1&lt;/span&gt; Summary&lt;/h3&gt;
&lt;p&gt;The &lt;em&gt;&lt;u&gt;connection between Gamma and Beta distribution&lt;/u&gt;&lt;/em&gt; helps us to find the normalization constant in Beta. In summary,&lt;/p&gt;
&lt;p&gt;If &lt;span class=&#34;math inline&#34;&gt;\(X \sim Gamma(\alpha, \lambda)\)&lt;/span&gt; and &lt;span class=&#34;math inline&#34;&gt;\(Y \sim Gamma(\beta, \lambda)\)&lt;/span&gt; are independent, then &lt;span class=&#34;math inline&#34;&gt;\(\frac{X}{X+Y} \sim Beta(\alpha, \beta)\)&lt;/span&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;plots&#34; class=&#34;section level2&#34; number=&#34;2.2&#34;&gt;
&lt;h2&gt;&lt;span class=&#34;header-section-number&#34;&gt;2.2&lt;/span&gt; plots&lt;/h2&gt;
&lt;pre class=&#34;r&#34;&gt;&lt;code&gt;library(zetaEDA)
library(ggfortify)
enable_zeta_ggplot_theme()&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Let’s check Beta density for some different parameters value.&lt;/p&gt;
&lt;p&gt;How about &lt;span class=&#34;math inline&#34;&gt;\(a = b = 1\)&lt;/span&gt;? The Beta(1,1) is just the Unif(0,1).&lt;/p&gt;
&lt;pre class=&#34;r&#34;&gt;&lt;code&gt;ggdistribution(func = dbeta, x = seq(0, 1, .01), shape1 = 1, shape2 = 1) +
  labs(title = &amp;quot;Beta Density with a = 1, b = 1&amp;quot;)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img src=&#34;https://chenxing.space/blog/note-for-beta-distribution/index.en_files/figure-html/unnamed-chunk-2-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;
&lt;p&gt;How about &lt;span class=&#34;math inline&#34;&gt;\(a = b = \frac{1}{2}\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://chenxing.space/blog/note-for-beta-distribution/index.en_files/figure-html/unnamed-chunk-3-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;
&lt;p&gt;How about &lt;span class=&#34;math inline&#34;&gt;\(a = b = 2\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://chenxing.space/blog/note-for-beta-distribution/index.en_files/figure-html/unnamed-chunk-4-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;
&lt;p&gt;How about &lt;span class=&#34;math inline&#34;&gt;\(a= 2, \ b = 1\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://chenxing.space/blog/note-for-beta-distribution/index.en_files/figure-html/unnamed-chunk-5-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;
&lt;p&gt;One more,&lt;/p&gt;
&lt;pre class=&#34;r&#34;&gt;&lt;code&gt;p &amp;lt;- ggdistribution(func = dbeta, x = seq(0, 1, .01), shape1 = 1.5, shape2 = 5, colour = &amp;quot;tomato&amp;quot;, linetype = &amp;quot;dashed&amp;quot;)

ggdistribution(func = dbeta, x = seq(0, 1, .01), shape1 = 5, shape2 = 1.5, colour = &amp;quot;blue&amp;quot;, p = p) +
  labs(title = &amp;quot;Red: a = 1.5, b = 5\n Blue: a = 5, b = 1.5&amp;quot;)&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img src=&#34;https://chenxing.space/blog/note-for-beta-distribution/index.en_files/figure-html/unnamed-chunk-6-1.png&#34; width=&#34;672&#34; /&gt;&lt;/p&gt;
&lt;p&gt;For more checking, click &lt;strong&gt;&lt;a href=&#34;https://homepage.divms.uiowa.edu/~mbognar/applets/beta.html&#34;&gt;this link&lt;/a&gt;&lt;/strong&gt; and try some parameters to check the density curve.&lt;/p&gt;
&lt;p&gt;Have fun!&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div id=&#34;reference&#34; class=&#34;section level1&#34; number=&#34;3&#34;&gt;
&lt;h1&gt;&lt;span class=&#34;header-section-number&#34;&gt;3&lt;/span&gt; Reference&lt;/h1&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;&lt;a href=&#34;https://en.wikipedia.org/wiki/Beta_distribution&#34;&gt;wiki Beta distribution&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href=&#34;https://stats.stackexchange.com/questions/47771/what-is-the-intuition-behind-beta-distribution&#34;&gt;What is the intuition behind beta distribution?&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=UZjlBQbV1KU&#34;&gt;youtube video Lecture 23: Beta distribution | Statistics 110&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href=&#34;https://www.youtube.com/watch?v=v1uUgTcInQk&#34;&gt;another youtube video: Beta distribution - an introduction&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Note for Gamma Distribution</title>
      <link>https://chenxing.space/blog/note-for-gamma-distribution/</link>
      <pubDate>Sat, 30 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/note-for-gamma-distribution/</guid>
      <description>

&lt;div id=&#34;TOC&#34;&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#motivation-for-gamma-function&#34;&gt;Motivation for Gamma Function&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#definition-of-gamma-function&#34;&gt;Definition of Gamma Function&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#gamma-distribution&#34;&gt;Gamma Distribution&lt;/a&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;#how-to-remember-the-gamma-pdf&#34;&gt;How to remember the Gamma pdf?&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;#gamma-exponential-connection&#34;&gt;Gamma &amp;amp; Exponential Connection&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/div&gt;

&lt;div id=&#34;motivation-for-gamma-function&#34; class=&#34;section level2&#34;&gt;
&lt;h2&gt;Motivation for Gamma Function&lt;/h2&gt;
&lt;p&gt;We all know how to compute the factorial of integer. BUT what is the factorial of 1/2?&lt;/p&gt;
&lt;p&gt;In other words, how to interpolate the factorial function?&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://tva1.sinaimg.cn/large/e6c9d24egy1h3r86q0n40j206y058a9z.jpg&#34; alt=&#34;img&#34; style=&#34;zoom:150%;&#34;/&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The gamma function can be seen as a solution to the following interpolation problem:&lt;/p&gt;
&lt;p&gt;“Find a smooth curve that connects the points (x, y) given by y = (x − 1)! at the positive integer values for x.”&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;More details, check the &lt;a href=&#34;https://en.wikipedia.org/wiki/Gamma_function&#34;&gt;wiki page&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;definition-of-gamma-function&#34; class=&#34;section level2&#34;&gt;
&lt;h2&gt;Definition of Gamma Function&lt;/h2&gt;
&lt;div class=&#34;definition&#34;&gt;
&lt;p&gt;&lt;span id=&#34;def:unnamed-chunk-2&#34; class=&#34;definition&#34;&gt;&lt;strong&gt;Definition 1  &lt;/strong&gt;&lt;/span&gt;(Gamma Function)
&lt;span class=&#34;math display&#34;&gt;\[\begin{align}
\Gamma(z) = \int_{0}^{\infty}x^{z-1}e^{-x}dx, \ \ z \in \mathbb{R}^+
\end{align}\]&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;For the the &lt;strong&gt;Gamma&lt;/strong&gt; &lt;strong&gt;function&lt;/strong&gt;, it is enough to know the following properties for now.&lt;/p&gt;
&lt;div class=&#34;lemma&#34;&gt;
&lt;p&gt;&lt;span id=&#34;lem:unnamed-chunk-3&#34; class=&#34;lemma&#34;&gt;&lt;strong&gt;Lemma 1  &lt;/strong&gt;&lt;/span&gt;&lt;span class=&#34;math display&#34;&gt;\[\begin{align}
&amp;amp; \Gamma(z+1) = z\Gamma(z), \ \ z \in \mathbb{R}^+ \\
&amp;amp; \Gamma(n) = (n-1)!, \  \ n = 1,2,3,...
\end{align}\]&lt;/span&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Easy to prove using integration by parts.&lt;/p&gt;
&lt;p&gt;Now, what is &lt;span class=&#34;math inline&#34;&gt;\(\Gamma(\frac{1}{2})\)&lt;/span&gt;?&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\Gamma(\frac{1}{2}) = \int_{0}^{\infty}x^{-1/2}e^{-x}dx = ?
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Recall that, &lt;span class=&#34;math inline&#34;&gt;\(\int_{0}^{\infty} e^{-x^2} = \frac{1}{2}\sqrt{\pi}\)&lt;/span&gt;, let &lt;span class=&#34;math inline&#34;&gt;\(u = x^2\)&lt;/span&gt; and will get the result &lt;span class=&#34;math inline&#34;&gt;\(\Gamma(\frac{1}{2} )= \sqrt{\pi}\)&lt;/span&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;div id=&#34;gamma-distribution&#34; class=&#34;section level2&#34;&gt;
&lt;h2&gt;Gamma Distribution&lt;/h2&gt;
&lt;p&gt;From the Gamma function, it is pretty natural to get Gamma pdf. JUST &lt;strong&gt;normalizing&lt;/strong&gt;!&lt;/p&gt;
&lt;p&gt;Clearly,
&lt;span class=&#34;math display&#34;&gt;\[
1 = \int_0^\infty \frac{x^{r-1}e^{-x}}{\Gamma(r)}dx = \int_0^\infty f_X(x)dx, \ \ \ X := Gamma(r, 1)
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;What is the pdf for the general &lt;span class=&#34;math inline&#34;&gt;\(Gamma(r, \lambda)\)&lt;/span&gt; ? Let&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
Y = \frac{X}{\lambda}, \ Y \sim Gamma(r, \lambda)
\]&lt;/span&gt;
&lt;span class=&#34;math display&#34;&gt;\[
f_Y(y) = f_X(x)\frac{dx}{dy}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;We’ll get&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
f(y; r, \lambda ) = \frac{\lambda ^{r}y^{r-1}e^{-\lambda y}}{\Gamma(r)}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Here, &lt;span class=&#34;math inline&#34;&gt;\(r\)&lt;/span&gt; is called the &lt;strong&gt;shape&lt;/strong&gt; parameter and &lt;span class=&#34;math inline&#34;&gt;\(\lambda\)&lt;/span&gt; is called the &lt;strong&gt;rate&lt;/strong&gt; parameter.&lt;/p&gt;
&lt;div id=&#34;how-to-remember-the-gamma-pdf&#34; class=&#34;section level3&#34;&gt;
&lt;h3&gt;How to remember the Gamma pdf?&lt;/h3&gt;
&lt;p&gt;That’s my trick: exponential density times the power rise to (shape-1), then divided by normalizer.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Exponential&lt;/strong&gt; density (very familiar): &lt;span class=&#34;math inline&#34;&gt;\(\lambda e^{-\lambda x}\)&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;power&lt;/strong&gt; rise to (shape-1): &lt;span class=&#34;math inline&#34;&gt;\((\lambda x)^{r-1}\)&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;normalizing&lt;/strong&gt; constant (using shape): &lt;span class=&#34;math inline&#34;&gt;\(\Gamma(r)\)&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\begin{align}
f(x; r, \lambda ) &amp;amp;= \frac{\text{exp density} \cdot \text{power}^\text{shape-1} }{normalizer} \\
&amp;amp; = \frac{\lambda e^{-\lambda x}(\lambda x)^{r-1}}{\Gamma(r)} \\
&amp;amp; = \frac{\lambda ^{r}x^{r-1}e^{-\lambda x}}{\Gamma(r)}
\end{align}
\]&lt;/span&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;p&gt;Another way to remember is this:&lt;/p&gt;
&lt;ol style=&#34;list-style-type: decimal&#34;&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Exponential&lt;/strong&gt; key part:
&lt;span class=&#34;math display&#34;&gt;\[
e^{-\lambda x}
\]&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Add &lt;strong&gt;Power&lt;/strong&gt; part:
&lt;span class=&#34;math display&#34;&gt;\[
\lambda^{\square} x^{\square} e^{- \lambda x}
\]&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;multiply the &lt;em&gt;power part&lt;/em&gt; in exponential&lt;/li&gt;
&lt;li&gt;&lt;em&gt;&lt;u&gt;rate&lt;/u&gt;&lt;/em&gt; rises to &lt;em&gt;&lt;u&gt;shape&lt;/u&gt;&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;&lt;u&gt;variable&lt;/u&gt;&lt;/em&gt; rises to &lt;em&gt;&lt;u&gt;shape-1&lt;/u&gt;&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\lambda^{r} x^{r-1} e^{- \lambda x}
\]&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Add &lt;strong&gt;Normalizing&lt;/strong&gt; part:
&lt;span class=&#34;math display&#34;&gt;\[
\frac{\lambda^{r} x^{r-1} e^{- \lambda x}}{\Gamma(r)}
\]&lt;/span&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
&lt;div id=&#34;gamma-exponential-connection&#34; class=&#34;section level3&#34;&gt;
&lt;h3&gt;Gamma &amp;amp; Exponential Connection&lt;/h3&gt;
&lt;p&gt;Let’s recall the Poisson Process,
&lt;span class=&#34;math display&#34;&gt;\[
N_t = \text{number of arrials up to time t} \sim Pois(\lambda t)
\]&lt;/span&gt;
The number of arrivals in the &lt;strong&gt;disjoint&lt;/strong&gt; intervals are &lt;strong&gt;independent&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Let &lt;span class=&#34;math inline&#34;&gt;\(T_1\)&lt;/span&gt; be the time of &lt;strong&gt;1st&lt;/strong&gt; arrival,
&lt;span class=&#34;math display&#34;&gt;\[
P(T_1 &amp;gt; t) = P(N_t = 0) = e^{-\lambda t} \ \implies T_1 \sim Exp(\lambda)
\]&lt;/span&gt;
Now, let &lt;span class=&#34;math inline&#34;&gt;\(T_n\)&lt;/span&gt; be the time of &lt;strong&gt;nth&lt;/strong&gt; arrival, that is, &lt;span class=&#34;math inline&#34;&gt;\(T_n = \sum_{i=1}^n X_i\)&lt;/span&gt; , where &lt;span class=&#34;math inline&#34;&gt;\(X_i \overset{\text{iid}}{\sim}Exp(\lambda)\)&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;What is the pdf of &lt;span class=&#34;math inline&#34;&gt;\(T_n\)&lt;/span&gt;? Answer is &lt;strong&gt;Gamma&lt;/strong&gt;!&lt;/p&gt;
&lt;div class=&#34;proposition&#34;&gt;
&lt;p&gt;&lt;span id=&#34;prp:unnamed-chunk-4&#34; class=&#34;proposition&#34;&gt;&lt;strong&gt;Proposition 1  &lt;/strong&gt;&lt;/span&gt;Gamma is the sum of iid Exponentials.&lt;/p&gt;
&lt;/div&gt;
&lt;p&gt;Proof:&lt;/p&gt;
&lt;p&gt;Since &lt;span class=&#34;math inline&#34;&gt;\(M_X(t) = \frac{\lambda}{\lambda-t}\)&lt;/span&gt;, where &lt;span class=&#34;math inline&#34;&gt;\(t &amp;lt; \lambda\)&lt;/span&gt;,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
M_{\sum_{i=1}^n X_i}(t) = (M_X(t))^n = (\frac{\lambda}{\lambda-t})^n
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;It is enough to show the MGF of Gamma equals the above value.&lt;/p&gt;
&lt;p&gt;Let &lt;span class=&#34;math inline&#34;&gt;\(Y \sim Gamma(n, \lambda)\)&lt;/span&gt;,&lt;/p&gt;
&lt;p&gt;&lt;span class=&#34;math display&#34;&gt;\[
\begin{align}
M_Y(t) = E(e^{ty}) &amp;amp;= \int_0^\infty e^{ty} \frac{1}{\Gamma(n)} \lambda^{n}e^{-\lambda y}y^{n-1} dy \\
&amp;amp;= \frac{\lambda^n}{\Gamma(n)} \int_0^\infty e^{-(\lambda - t)y}y^{n-1}dy, \text{ let } u =  (\lambda - t)y \\
&amp;amp;= \frac{\lambda^n}{\Gamma(n)} (\frac{1}{\lambda-t})^n \int_0^\infty e^{-u}u^{n-1}du \\
&amp;amp;= \frac{\lambda^n}{\Gamma(n)} (\frac{1}{\lambda-t})^n\Gamma(n) \\
&amp;amp;= (\frac{\lambda}{\lambda-t})^n
\end{align}
\]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Proved!&lt;/p&gt;
&lt;div class=&#34;remark&#34;&gt;
&lt;p&gt;&lt;span id=&#34;unlabeled-div-1&#34; class=&#34;remark&#34;&gt;&lt;em&gt;Remark&lt;/em&gt;. &lt;/span&gt;The &lt;strong&gt;Exponential&lt;/strong&gt; is the continuous analog of the &lt;strong&gt;Geometric&lt;/strong&gt;. Similarly, the &lt;strong&gt;Gamma&lt;/strong&gt; is the continuous analog of the &lt;strong&gt;Negative Binomial&lt;/strong&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
</description>
    </item>
    
    <item>
      <title>Deriving the conditional distributions of a multivariate normal distribution</title>
      <link>https://chenxing.space/blog/deriving-the-conditional-distributions-of-a-multivariate-normal-distribution/</link>
      <pubDate>Fri, 31 Mar 2017 00:00:00 +0000</pubDate>
      <guid>https://chenxing.space/blog/deriving-the-conditional-distributions-of-a-multivariate-normal-distribution/</guid>
      <description>&lt;p&gt;Overall, the intuition behind the conditional distribution of a bivariate normal is that even though the two variables are correlated in the joint distribution, they can still be treated as independent when you&amp;rsquo;re looking at the distribution of one variable, given a fixed value of the other variable. This allows you to use the normal distribution, which is a well-understood and widely-used probability distribution, to model the conditional distribution of a bivariate normal.&lt;/p&gt;
&lt;p&gt;More generally,&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202303311241754.png&#34; alt=&#34;image-20230331124124725&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;Assume we know there&amp;rsquo;s a theorem that says all conditional distributions of a multivariate normal distribution are normal.&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img src=&#34;https://raw.githubusercontent.com/chenx2018/blogdown-image/main/img/202303311238287.png&#34; alt=&#34;image-20230331123847757&#34; loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
