In the first part I showed that the SAXS intensity scattered by a platelet system goes like \( I(q) \sim q^{-2}\), at least in some intermediate (but as yet unspecified) q range. Here I will show that for thin rods this dependence becomes \( q^{-1}\), I will then derive the terminal (Porod) behaviour \( q^{-4}\) and briefly consider the transition between these two regimes.
April 27, 2023
Power laws in small-angle scattering - part I
The small-angle X-ray scattering (SAXS) spectrum of particles with a well-defined shape (such as rods or platelets) is often characterized by a power-law dependence: \( I(q) \sim q^{-\alpha}\), where the exponent \( \alpha \) is directly related to the particle geometry. For "compact" particles, the large-\( q \) intensity scales as \( q^{-4}\) (Porod regime). Below, I'll give the most compact and yet -hopefully- understandable derivation I can think of for these power laws.
To simplify the derivation, we'll consider these objects as infinitely thin and infinitely large, meaning that we'll be looking at them on length scales much larger than their thickness and much smaller than their lateral extension. The approximation is legitimate, since it is in this range of length (or, conversely, scattering vector) that the power-law regimes are encountered.
As discussed above, the Patterson function is similar to the density and thus we will apply the same approximation to \(P(\mathbf{r})\), which is the natural descriptor of the system, due to its intimate relation with the intensity \(I(\mathbf{q}) = \left | \tilde{\rho}(\mathbf{q}) \right |^2\).
August 29, 2021
The Pareto distribution and Price's law
As detailed in the previous post, the ratio \(f\) of the top authors that publish a fraction \(v\) of all publications is independent from the total number of authors \(N_0\). Of course, this result is incompatible with Price's law (that for \(v=0.5\), \(f = 1/\sqrt{N_0}\)). This issue has been discussed by Price and co-workers [1], but I will take here a slightly different approach.
I had assumed in my derivation that he domain of the distribution was unbound above (\(H = \infty\)), and that the exponent \(\alpha\) was higher than 1. One can relax these assumptions and check their effect on \(f\) by:
- imposing a finite upper bound \(H\) and
- by setting \(\alpha = 1\). Note that 2. also requires 1.
Role of the upper bound
In the finite \(H\) case one must use the full expressions (containing \(H\) and \(L\)) for the various quantities. In this section, we will continue to assume that \(\alpha > 1\). Since \(L\) acts everywhere as a scale factor for \(x\) (and \(H\)) I will set it to 1 in the following. It is also reasonable to assume that the least productive authors have one publication (why truncate at a higher value?!) Consequently, all results will also depend on \(H\), but presumably not explicitly on \(N_0\), which is a prefactor for the PDF and should cancel out of all expectation calculations. It is, however, quite likely that \(H\) itself will depend on \(N_0\), since more authors will lead to a higher maximum publication number!
In my opinion, the most reasonable assumption is that there is only one author with \(H\) publications, so that \(N_0 p(H) = 1 \Rightarrow H \simeq (N_0 \alpha)^{\frac{1}{\alpha + 1}}\), neglecting the normalization prefactor of \(p(x)\).
The threshold number \(x_f\) is easy to obtain directly from \(S(x)\):
\[x_f = \left [ f + (1-f) H^{-\alpha}\right ]^{-1/\alpha}\]
From its definition, the fraction \(v\) is given by: \(v = \dfrac{\alpha}{\mu} \dfrac{1}{1-H^{-\alpha}} \dfrac{1}{\alpha - 1} \left ( x_f^{1-\alpha} - H^{1-\alpha} \right )\). Note that we need here the complete expression for the mean [2]:
\[\mu = \dfrac{\alpha}{\alpha - 1} L \dfrac{1-H^{1-\alpha}}{1-H^{-\alpha}}\]
Plugging \(x_f\) and \(\mu\) in the definition of \(v\) and setting \(v = 1/2\) yields:
\begin{equation} f = f_{\infty} \dfrac{\left ( 1 + H^{1-\alpha}\right )^{\frac{\alpha}{\alpha - 1}} - 2^{\frac{\alpha}{\alpha - 1}}H^{-\alpha}}{1-H^{-\alpha}}, \quad \text{with } f_{\infty} = \left( \dfrac{1}{2} \right )^{\frac{\alpha}{\alpha - 1}},\end{equation}and we assume that the upper bound is given by:
\begin{equation} H = (N_0 \alpha)^{\frac{1}{\alpha + 1}}. \end{equation}Exponent \(\alpha = 1\)
Let us rewrite the PDF, CDF and survival function in this particular case:
\[p(x) = \dfrac{1}{1 - H^{-1}} \dfrac{1}{x^2}; \, F(x) = \dfrac{1- x^{-1}}{1 - H^{-1}} ; \, S(x) = 1 - F(x) = \dfrac{x^{-1}- H^{-1}}{1 - H^{-1}}\]
\[x_f = S^{-1}(f) = \dfrac{1}{f + (1-f) H^{-1}}\]
\[v = \dfrac{1}{2} = 1 - \dfrac{\ln(x_f)}{\ln(H)} \Rightarrow x_f = \sqrt{H} \quad \text{and, since } H = \sqrt{N_0}, \, x_f = N_0^{1/4}\]
Putting it all together yields \(f = \dfrac{N_0^{1/4} - 1}{N_0^{1/2} - 1}\) and, in the high \(N_0\) limit, \(f \sim N_0^{-1/4}\), so the number of "prolific" authors \(N_p = f N_0 = N_0^{3/4}\), a result also obtained by Price et al. [1] using the discrete distribution. They also showed that other power laws (from \(N_0^{1/2}\) to \(N_0^{1}\)) can be obtained, depending on the exact dependence of \(H\) on \(N_0\).
1 Allison, P. D. et al., Lotka's Law: A Problem in Its Interpretation and Application Social Studies of Science 6, 269-276, (1976).↩
August 28, 2021
The Pareto distribution and the 20/80 rule
I mentioned in the previous post Pareto's 20/80 rule. Here, I will discuss Pareto's distribution, insisting on how (and in what conditions) it gives rise to this result. I had some trouble understanding the derivation as presented in various sources, so I will go through it in detail.
The functional form of the Pareto distribution is a power law, over an interval \((L,H)\) such that \(0<L<H\leq \infty\). I will use the notations of the Wikipedia page unless stated otherwise. Its probability density function (PDF) \(p(x)\) and cumulative distribution function (CDF) \(F(x)\) are (\(\alpha\) is real and strictly positive):
\[p(x) = \dfrac{\alpha}{1-(L/H)^{\alpha}} \dfrac{1}{x} \left ( \dfrac{L}{x} \right ) ^{\alpha}\quad ; \quad F(x) = \dfrac{1-(L/x)^{\alpha}}{1-(L/H)^{\alpha}}\]
One often uses the complementary CDF (or survival function) defined as:
\[S(x) = 1 - F(x) = \dfrac{1}{1-(L/H)^{\alpha}}\left [ \left ( \dfrac{L}{x} \right )^{\alpha} - \left ( \dfrac{L}{H} \right )^{\alpha}\right ]\]
Note that the survival function is very similar to the PDF multiplied by \(x\): \(S(x) \simeq \dfrac{x}{\alpha} p(x)\), the difference being due only to the final truncation term. However, this is only true for power laws, as one can easily check by writing \(p(x) = F'(x)\) and solving the resulting ODE. We should therefore carefully distinguish \(x p(x)\) (which is, for instance, the integrand to use for computing the mean of the distribution) and \(S(x)\) which "has already been integrated", so to speak.
Let us use this continuous model to describe the distribution of publications (neglecting for now its intrinsically discrete character). \(x\) stands for the number of publications by one author, bounded by \(L\) and \(H\). The number of authors that published \(x\) books is given by \(N_0 \, p(x)\). \(N_0\) is the total number of authors.
- The first question is: who are the first \(f\) more prolific authors (in Pareto's case, \(f = 0.2 = 20\)%)? More precisely, what is the threshold number of publications \(x_f\) separating them from the less prolific ones?
- The second question is: how many publications did these top \(f\) authors contribute?
