| ||||||||
About the bulge in the variance, that one can see in 04wLu or 03BPi or 03avx, this is not due to merely increasing the combinatorial size of the sampling, since the max remains at fairly small $k$ (it should be at $k\approx N/2$ if due to $\binom Nk$). Instead, it comes from the leading term of: \begin{equation} \label{eq:05hUD} \operatorname{Var}_\mathcal{C}(p_\mathcal{C})=\sum_{j=1}^{k}\frac{\binom kj^2\,\delta_j(k)}{\binom Nj} =\frac{k^2\zeta_1(k)}{N}+\frac{k^2(k-1)^2\,\zeta_2(k)}{2N^2}+O(N^{-3})\,. \end{equation}
This is $\frac{k^2\zeta_1}{N}$ for dipole-like states, and this shows the bulge: $k^2$ increases with $k$ while $\zeta_1$ decreases. The physical fight goes as follows:
- $k^2$'s increase comes from the fluctuation of the collapse mean being the average of the $k$ vertices' contribution which goes as $(k/N)\sum_i g_1(\theta_i)$, so each extra vertex adds one more channel by which the collapse's photons can move the answer.
- $\zeta_1(k)$'s decrease comes from the finer polygon worrying less about each additional single vertex.
The turnover occurs when $\zeta_1$ decreases faster than $k^{-2}$. This turns out to be robustly around $k \approx10$.
This was for dipole-like. For the donut, $\zeta_1=0$ and one needs to do this analysis with the $k^4$ dependence. Then the buldge is at smaller $k\approx4$.
More importantly, however, this might not be quite relevant as the gain itself (04bSr) is what matters most, and this is monotonously decreasing. And what's more, $\operatorname{Var}_\mathcal{C}(p_\mathcal{C})$ doesn't even appear to be the most important factor, as compared to the variance $\sigma^2$ itself. So not much to dig in that direction. I'll leave it there.
In the previous 𝕫eet, the behaviour of $\zeta$ as a variance gets to look counter-intuitive. It increases with the number of added points while one could expect it to decrease. It's actually a good way to better understand what this is.
Once $c$ vertices are known, the perimeter still has a remaining spread, $\operatorname{Var}(p\mid c\text{ vertices})$. This one does what intuition commands: it decreases with $c$ (more information, more pinpointed), from $\sigma^2$ at $c=0$ down to $0$ at $c=k$.
But $\zeta_c$ is not that. $\zeta_c$ is the spread of the prediction, $\operatorname{Var}\bigl(E[p\mid c\text{ vertices}]\bigr)$: how much the best guess differs from one telling to another. This actually increases with $c$: told nothing, the guess is always $\langle p\rangle$ (spread $0$); told everything, the guess is the perimeter itself (spread $\sigma^2$).
The two are complementary, by Eve's law, exactly as within/between for the collapse: \begin{equation} \sigma^2=\underbrace{\operatorname{Var}\bigl(E[p\mid c\text{ vertices}]\bigr)}_{\zeta_c}+\underbrace{E\bigl[\operatorname{Var}(p\mid c\text{ vertices})\bigr]}_{\text{residual}_c}\,. \end{equation} The sum is fixed at $\sigma^2$; as $c$ grows, the explained part $\zeta_c$ goes up and the residual goes down by the same amount. So $\zeta_c/\sigma^2$ is the share of the perimeter's variance that $c$ vertices explain, the same role that $\rho$ plays for the collapse ($\operatorname{Var}_\mathcal{C}(p_\mathcal{C})$ is what the collapse explains, $E[v_\mathcal{C}]$ the residual).
In short: $\zeta_c$ is the part of $\sigma^2$ that $c$ vertices explain; the remaining uncertainty $\sigma^2-\zeta_c$ shrinks as more vertices are known.
I now relate our findings to U-statistics, since this is precisely an instance of that, that we rediscovered independently.
$p_\mathcal{C}$ is a $U$-statistic. This allows us to derive the following results (derivations and comments on ℤ): \begin{equation} \operatorname{Var}_\mathcal{C}(p_\mathcal{C})=\sum_{j=1}^{k}\frac{\binom kj^2\,\delta_j}{\binom Nj} % =\frac{k^2\zeta_1}{N}+\frac{k^2(k-1)^2\,\delta_2}{2N(N-1)}+O(N^{-3}), \label{eq:04zq8} \end{equation} while the independent-ensemble variance is Eq. \eqref{eq:04zq8} at $k=N$ (a collapse of $k$ photons is one $k$-gon), which gives: \begin{equation}\label{eq:04UFj}\sigma^2=\sum_{j=1}^{k}\binom kj\delta_j\,.\end{equation}
In the above, $\delta_j\equiv\operatorname{Var}(g_j)$ where $g_j$ are terms that appear in the Hoeffding decomposition, which I detail on ℤ.
The perimeter $p(\theta_1,\dots,\theta_k)$ depends on $k$ random vertices. Pin $c$ of them at fixed angles $\vartheta_1,\dots,\vartheta_c$ and average over the other $k-c$, which are photons drawn from $f$, to obtain the $c$th function $h_c$: \begin{gather*} h_0=\langle p\rangle,\\ h_1(\vartheta_1)=E\bigl[p\,\big|\,\text{one vertex at }\vartheta_1\bigr],\\ h_2(\vartheta_1,\vartheta_2)=E\bigl[p\,\big|\,\text{vertices at }\vartheta_1,\vartheta_2\bigr],\\ \cdots\\ h_k(\vartheta_1,\dots,\vartheta_k)=p(\vartheta_1,\dots,\vartheta_k)\,. \end{gather*} Each can be obtained from the next one by averaging over the extra photon, which is drawn from $f$: $\int h_1(\vartheta_1)f(\vartheta_1)\,d\vartheta_1=h_0$, $\int h_2(\vartheta_1,\vartheta_2)f(\vartheta_2)\,d\vartheta_2=h_1(\vartheta_1)$, etc.
The theory of $U$-statistics gives their mean—that is Adam's law—and their variance. The focus is on the variance, as the mean does what Adam's law wants it to do, which is simple. So we now study the variances of the $h_c$ when the pinned angles are themselves drawn from $f$: $$\zeta_c\equiv\operatorname{Var}\,h_c(\theta_1,\dots,\theta_c)\,,$$ with the following hierarchy: $$0=\zeta_0\le\zeta_1\le\dots\le\zeta_k=\sigma^2\,.$$ $\zeta_c$ measures how much knowing $c$ of the vertices moves the expected perimeter: not at all for $c=0$, entirely for $c=k$. That is to say, if to guess the perimeter of a $k$-gon whose photons have not landed yet, the best guess is $\langle p\rangle$, with variance $\zeta_0$. Told where one vertex is, $\vartheta_1$, the best guess now becomes $h_1(\vartheta_1)$—the mean perimeter of $k$-gons with a vertex there. And now the guess depends on this information: for some $\vartheta_1$, it is above $\langle p\rangle$, for others, it is below. $\zeta_1$ is the variance of that guess as $\vartheta_1$ varies (with the photon's law $f$). Told $c$ vertices, the guess is $h_c(\vartheta_1,\cdots,\vartheta_c)$ with variance $\zeta_c$. The more vertices, the more can the guess differ from $\langle p\rangle$, so $\zeta_c$ actually grows with $c$. Told all $k$ vertices, the guess is the fixed perimeter itself, which scatters with the full $\sigma^2$ of the ensemble, so $\zeta_k=\sigma^2.$
Now come the $g_j$. These are the components of $p$ that are centred in each of their arguments, so that they average to zero over any one of them, and this makes them uncorrelated: \begin{align*} g_1(\vartheta_1)&=h_1(\vartheta_1)-\langle p\rangle,\\ g_2(\vartheta_1,\vartheta_2)&=h_2(\vartheta_1,\vartheta_2)-h_1(\vartheta_1)-h_1(\vartheta_2)+\langle p\rangle,\\ g_3(\vartheta_1,\vartheta_2,\vartheta_3)&=h_3(\vartheta_1,\vartheta_2,\vartheta_3)-h_2(\vartheta_1,\vartheta_2)-h_2(\vartheta_1,\vartheta_3)-h_2(\vartheta_2,\vartheta_3)\\ &\qquad+h_1(\vartheta_1)+h_1(\vartheta_2)+h_1(\vartheta_3)-\langle p\rangle, \end{align*} and so on by inclusion–exclusion: $g_1$ is what one vertex does on its own, $g_2$ what a pair does beyond what each does alone, etc.
This apparently comes in a variety of statistical models under different trades and names, including the functional ANOVA or Sobol decomposition of sensitivity analysis, etc. We could spend not even the rest of the day but the rest of the year in studying all this.
Anyway, the really important thing is that each $g_j$ averages to zero over any one of its arguments. For instance, if we average $g_2(\vartheta_1,\theta_2)$ over the photon $\theta_2$ (drawn from $f$), then we find $h_1(\vartheta_1)-h_1(\vartheta_1)-\langle p\rangle+\langle p\rangle=0$. The same cancellation holds for every $g_j$ and every one of its arguments. So we can write: \begin{equation} \label{eq:04m5O} p(\theta_1,\dots,\theta_k)=\langle p\rangle+\sum_i g_1(\theta_i) +\sum_{i<j}g_2(\theta_i,\theta_j)+\dots+g_k(\theta_1,\dots,\theta_k) \end{equation} and any two distinct terms are uncorrelated: one always contains a photon the other does not, and averaging over that photon kills the term. As a consequence, the variances in Eq. \eqref{eq:04m5O} add. With $\delta_j\equiv\operatorname{Var}g_j$, then $\sigma^2=k\,\delta_1+\binom k2\delta_2+\dots+\delta_k$, which is Eq. \eqref{eq:04UFj}. The variance of independent sampling in this way splits into what single vertices, pairs, triples, etc., contribute.
The two types of variances relate to each other as follows: $$\zeta_c=\sum_{j\le c}\binom cj\delta_j\,,$$ i.e., $\zeta_1=\delta_1$, $\zeta_2=\delta_2+2\delta_1$, etc.
From their above definition, $\zeta_c$ is the variance of a conditional mean, or equally the covariance of two $k$-gons sharing exactly $c$ vertices (the others independent). It is what one measures from the data. On the other hand, $\delta_j$ is the variance of the pure $j$-body component, with everything of lower order subtracted, and it is what enters the formulas cleanly.
Now, the mean over one collapse, i.e., $p_\mathcal{C}$, is the average of Eq. \eqref{eq:04m5O} over all the $\binom Nk$ subsets $S$ of the $N$ photons. In this average, by symmetry, every one of the $\binom Nj$ $j$-subsets of the collapse receives the same weight $w_j$ in front of its $g_j$. Each copy of Eq. \eqref{eq:04m5O} contains $\binom kj$ terms of order $j$ with unit weight, and averaging preserves the total weight, so $$w_j=\binom kj\Big/\binom Nj\,.$$ (Equivalently: a given $j$-subset lies in $\binom{N-j}{k-j}$ of the $k$-subsets $S$, and $\binom{N-j}{k-j}/\binom Nk=\binom kj/\binom Nj$.) The $\binom Nj$ terms of order $j$ are still uncorrelated, each with variance $\delta_j$, so their variances add: the order-$j$ part of $p_\mathcal{C}$ has variance $w_j^2\times\binom Nj\times\delta_j=\binom kj^2\delta_j\big/\binom Nj$—this is where the square comes from—and summing the orders: $$ \operatorname{Var}_{\mathcal{C}}(p_{\mathcal{C}}) =\sum_{j=1}^{k}\frac{\binom kj^2\delta_j}{\binom Nj},$$ which is Eq. \eqref{eq:04zq8}.
Importantly, the single-collapse must be exhaustive: we must take everything, the whole combinatorics lot. This is possible for particular types of observables thanks to linearity and additivity, but not in general.
If the collapse is instead read through $B$ random $k$-subsets, the estimate carries the within-collapse noise and by Eve's law: $$\operatorname{Var}_\mathcal{C}(p_\mathcal{C})+\frac{E[v_\mathcal{C}]}{B} =\sigma^2\Bigl[\rho+\frac{1-\rho}{B}\Bigr],$$ which beats independent sampling, $\sigma ^2k/N$, only for $B>(1-\rho)/({k\over N}-\rho)$. I remind what $\ρho$ is in the ℤ supplementary below. For $\rho\ll k/N$, that becomes $B>N/k$.
Reading $N/k$ random $k$-gons out of one collapse is therefore not better than $N/k$ fresh ones. It is, in fact, slightly worse, by the factor $1+\rho({N\over k}-1)$ in variance: $1.006$ at the donut, $1.23$ at the atom, $1.27$ at criticality for $k=10$.
Also a reminder on what $\rho$ is: at fixed $k$ and $N$, \begin{equation} \rho\equiv\frac{\operatorname{Var}_{\mathcal{C}}(p_{\mathcal{C}})}{\sigma^2} =1-\frac{E[v_{\mathcal{C}}]}{\sigma^2}\,. \label{eq:rhodef} \end{equation} It is the ratio of the widths of the single-collapse distribution vs that of the independent ensemble distribution.
$\rho\ll 1$ at small $k$ (the collapse mean is far steadier than one shot) and $\rho=1$ exactly at $k=N$. When it is 1/2, the noise is equally distributed.
It is related to the gain per photon $G$ as $G=k/(N\rho)$.
It should be computed from the distributions, rather than from the formulas above, that are numerically unstable at small $k$ given the huge imbalances between the variances.
We now compare both approaches in terms of how efficient they are as estimators: single-collapse of $N$ points vs $N'$ independent samplings of $k$ points to compute a $k$-photon observable.
Main result is, single-collapse is always at least as good, $G\ge 1$, where $G\equiv$ is the ratio of the two variances: \begin{equation} G(k,N)=\frac{\sigma^2/N'}{\operatorname{Var}_\mathcal{C}(p_\mathcal{C})} =\frac{k}{N}\,\frac{\sigma^2}{\operatorname{Var}_\mathcal{C}(p_\mathcal{C})} =\frac{k}{N\rho}\,, \label{eq:04SYs} \end{equation}
The reason is not that a collapse holds more photons ($N=kN'$), but that they are recombined into all $\binom Nk$ $k$-gons instead of $N/k$. The recombined $k$-gons overlap and are not independent—so one could assume bias or fake structure—but single-collapse estimator are unbiased by Adam's theorem and although the gain is very far from $\binom Nk/(N/k)$, it is still substantial (it is $G$).
That's still $G=500$ for the donut at $k=3$. The gain is less qualitative at or near criticality, since we have only $G\approx 2$ over a long plateau, but is still always a gain, even at $k=N-1$, where $G=1.0028$ (donut), $G=1.0019$ (atom) and $G=1.0008$ (dipole) (additional ℤ material below). Not much, but still $>1$.
All of them converge to $G=1$ exactly when $k=N$, since the two procedures reduce to the same one.
One can compute the particular case $G(N-1,N)$ as: \begin{equation}\label{eq:04rE9}G(N-1,N)=\frac{(N-1)\,\zeta_{N-1}}{\zeta_{N-1}+(N-1)\,\zeta_{N-2}},\end{equation} with $\zeta_{N-1}=\sigma^2(N-1)$ and $\zeta_{N-2}$ the covariance of two $(N-1)$-gons sharing $N-2$ vertices. These need be computed (numerically in most cases). This is strongly state dependent. We limit to the donut (most particular case but also the most interesting as the deviation is the biggest) which approximates to: $$G(N-1,N)=1+\frac{14}{5N}+O(N^{-2})\qquad(C_\alpha=0).$$ This illustrates the type of information one can get on general statements such as single-collapse $U$-statistics outperforming independent collapses.
Back to Eve's law, an explicit computation yields:
and lots of things can be seen here.
First, red $\ll$ blue until high $k$. A single collapse is considerably better than independent collapses. After a perfunctory look at it, that's indeed a manifestation of U-statistics. The latter seems particularly fitted for bosonic observables. It is so good that the single-collapse noise is essentially negligible as compared to normal statistical noise.
Red (single-collapse noise) is non-monotonous: it increases until low $k\approx 10$, and then decays.
Black (variance of ensembles) is strongly state dependence for high $k$. While $\sigma^2$ falls as $k^{-2.4}$ between $k=3$ and $10$ for all three states, it then decays sharply as $k^{-4.9}$ at the donut and $k^{-4.3}$ at the atom, but only as $k^{-2.0}$ at the dipole, from $k=10$ all the way to $700$. The critical state's single-shot variance decays two powers of $k$ slower than the donut's!
Adam's law—or law of total expectation—states that: \begin{equation}\label{eq:03QDi}E_\mathcal{C}\bigl[p_\mathcal{C}(k)\bigr]=\langle p\rangle (k)\end{equation} exactly, for every $k$.
Eve's law—or law of total variance—states that: \begin{equation} \underbrace{\Var(p)}_{\sigma^2} = \underbrace{E_\mathcal{C}\bigl[\Var(p\,|\,\mathcal{C})\bigr]}_{E[v_\mathcal{C}]} + \underbrace{\Var_\mathcal{C}\left(E[p|\mathcal{C}]\right)}_{\Var_\mathcal{C}{(p_\mathcal{C})}}, \label{eq:03zYK} \end{equation}
The variances involved are distributed over two types of samplings:
- The collapse. The state is measured: $N$ photons land on the circle at angles drawn iid from $f_{C_\alpha}(\theta)$. Output: the frozen set $\mathcal C=\{\theta_i\}_{i=1}^N$.
- A subset of the collapse. From that frozen $\mathcal C$, a $k$-subset is drawn uniformly among the $\binom{N}{k}$ of them, and its perimeter $p$ is read off.
For a given $k$:
definition averaged over is the width of independent samplings (ensemble average) sampling from within one collapse repeated collapses The important thing for us is that: $$\Var_\mathcal{C}(p_\mathcal{C})=\Var(p)-{E}_\mathcal{C}\left[\Var(p|\mathcal{C})\right]$$
This is meant for a conceptual reading, not a recipe to compute them, since for small $k$, the lhs vanishes in a way that is not numerically stable. For instance, for $N=1000$ data, at $C_\alpha=0$, $k=3$, we find $\Var(p)=1.135$ while $\Var_\mathcal{C}(p_\mathcal{C})=6.81\times10^{-6}$.
This, however, becomes useful numerically near $k=N$, where the two terms are comparable. Everywhere else $\Var_\mathcal{C}(p_\mathcal{C})$ should be computed directly from $p_\mathcal{C}$ over repeated collapses.
The conceptual reading is that while Adam's law Eq. \eqref{eq:03QDi} equates ensemble and collapse averages, Eve's law shows that the collapse variance is smaller—typically much smaller—than the ensemble average. What narrows it is how the variate fluctuates in a frozen collapse.
Therefore, if you take a single collapse, its one $p_\mathcal{C}$ will not be $\langle p\rangle$—you bet, you made one measurement—but it will be much closer to $\langle p\rangle$ than an average over independent ensemble would be!
That is definitely something that is known by someone, somewhere. But it's beautiful in our context. It needs to be formulated in terms of estimators. How much better the collapse $p_\mathcal{C}$ performs over the standard statistical estimator $\hat p$. Note that one can also draw several $p_\mathcal{C}$ and this becomes something like $\hat p_\mathcal{C}$ of order $i$ (if sampling $i$ of them, the beautiful thing is that even $i=1$ is a better estimator than the statistical one).
I'm very proud of Eq. \eqref{eq:03zYK} (Eq. 03zYK (af)), which is coming next post (equation in the future). It is very symmetric, swapping averages and variances over different statistical ensembles. And it's named after she!
J doesn't reproduce my results though. He gets kinkiness in the exact expression. But it's highly suspicious: the worst case is $k=7$ which affects all three collapses. Why? Noise, if any, should come from the collapses, nothing should happen for everybody at $k=7$. In fact if you would fix this one, it'd look like quite smooth (not exactly, though). Besides I've checked the regularity of the weighting terms even down to small $N$ and it is always smooth, baring some border discontinuities for few cases. So I think he's not summing everything.
The aesthetic of noise.
Or order and chaos. One vs many.
An idea of the ruggedness of the states $C_\alpha$ can be shown by spanning over all states $0\le C_\alpha\le 1/4$ for all $k$.
I show two cases so that one can see the variation from individual to individual. Like us in the streets, they both look like any other: 1m70, two eyes, X chromosomes. They only differ in details: one has a mole by the nose, the other has a faint strabismus, one has a Y chromosome, the other two X, etc. they differ in character too. The science of individuality from the laws of the group.
So as to show a smooth evolution, I sample $N=10^3$ points $u_i$ uniformly, and compute their inverse-transform sampling: $$\theta_i^{(C_\alpha)} = F_{C_\alpha}^{-1}(u_i),\qquad F_C(\theta)=\frac{\theta}{2\pi}+\frac{V\sin2\theta}{4\pi},\quad V=2\sqrt{C_\alpha},$$ in this way, we obtain this beautiful smooth transition from donut (0) till criticality (1/4). Each state belongs to the $C_\alpha$ family but is connected to its neighbours so that the structure propagates.
The smoothness of the curves surprisingly does not even come from the fact that we are dealing with huge data reservoirs, e.g., with a single $N=1500$ collapse. That's already 0.6 billion points for perimeters of triangles. But for $N=10$, that's 252 data points max (for pentagons). That still makes a smooth curve:
This comes from the smoothness of the ${\binom{N-m-1}{k-2}\big/\binom{N}{k}}$ weight itself, which smooths out the real disordered, rugous data $G_{i,m}-2\sin\left(G_{i,m}/2\right)$ from the peculiarities of the collapse.
Now, if we don't take all the ${N\choose k}$ data points (as is possible thanks to the additivity property of such variables), then we break the smooth combinatorial balance and introduce noise again.
Our sampling of $n$ data points instead thus turns into an estimator $\hat p_\mathcal{C}$ of the sample mean, so its error is that given by standard statistics as $\sigma_\mathcal{C}/\sqrt n$ with $\sigma_\mathcal{C}$ the std of $D(p\,|\,\mathcal{C})$. These are a few cases of all-$k$ trajectories that show that, for different $n$ (how many samplings to build the $D(p|\mathcal{C})$ distribution from one collapse) and $N$ (how many total points for the collapse).
The individual vs the ensemble.
Dash-dotted red is everybody. Blue is one specimen. We follow them as we probe more and more of their structure until we look at their everything.
Not such a big deal, he?
This was for $C_\alpha=0$, the mootest, shapeless-looking subset, where each individual actually looks like the full set, until very close to seeing everything, when they finally thin down into their final, ultimate Dirac $\delta$ form: they become who they are.
If we look at $C_\alpha=1/6$, the atom for the thermal distribution, we get something already more interesting:
Some internal struggle, some wobbling, before the thinning out.
But wait before you see the critical case, $C_\alpha=1/4$. That's chaotic:
The first ℤ post was motivated by Misha Glazov's comments on Bosonization. He replied almost immediately. Me, as the rude, inconsiderate person I am, only reply to his reply now.
I'm not going to quote the entirely of a long and very sophisticated email, full of anecdotes, delicious comments such as «The devil is, of course, in the initial approximations» or quotes such as «On the other hand, I am not entirely happy with general statements that everything is conceptually wrong with excitons. It sounds a bit like “all papers are wrong but my last one” [(c) M.I. Dyakonov]», but stick to points of my own reply, so that I can keep track.
On the broader question, no clear-cut conclusion:
Quick answer would be: I cannot provide a simple and justified one-to-one correspondence between bosonization of Wannier-Mott excitons and bosonization of Luttinger liquid. Both approaches are approximate, the approximations are different. Whether this difference is purely technical or has a very deep physics inside, I do not know. Some ideas are below.From what I understand of your ideas below, it appears there is little to connect together and that the same (natural) term is used for two not related—or not so clearly related—things.
I guess the most basic way to reformulate my question in the light of your explications is: "Is bosonization limited to Luttinger liquids or can it extend to other classes of fermionic systems, that intrinsically exhibit bosonic properties (excitons, Cooper pairs, etc.)"
I've already been assured that the technique is not specific to 1D.
Second, maybe, the simplest start would be if you remind be what exactly is Combescot’s point in criticising excitons (at least how you understand it)? Then we can try to apply her arguments to bosonization of 1D systems and with a bit of effort and luck learn something. Otherwise, for me it is extremely hard to look for an answer to unknown question.Right, this is indeed a necessary starting point. The crux of her argument is that she does not criticize the failure of some approximation when pushed into a regime where it ceases to apply, e.g., at large densities, but that the concept is structurally wrong from the outset. The point is that the object itself—the exciton—cannot be defined because its interaction with the rest of the universe requires to trade its own internal constituents with it, e.g., an exciton-electron interaction makes the exciton's electron indistinguishable from the one outside of itself. So what is the exciton? What is not the exciton? This is ill-defined. Interestingly, this is not problem of a weak or strong effect (perturbative parameter): it is structural in the sense that even in absence of Coulomb interaction, the electron-electron exchange would result in some dynamics / interaction (but without interactions, no binding). This cannot be described by an effective Hamiltonian. This is my very simplified understanding of her main point.
Now, I don't know how valid and important this is, but I can feel it could be. Her objection also applies to 1D systems, where Haldane bosonization is exact. So in principle—my thinking is—one could compare three treatments of the problem:
- Haldane (exact)
- Combescot (validity unknown)
- Effective bosonic Hamiltonian (validity unknown)
And see if 2=1 or 3=1 or if neither 2 or 3 equal 1. Now you are saying that Luttinger liquids don't involve excitons, in the sense of single-particle excitations, only collective excitations. But I don't see this as settling my question because this could be precisely what Combescot is saying, that the exciton should not be dealt with as a single particle. You say it in the sense that other systems have well-defined single-particle excitations, like atomic-looking lines. But Koch also says that those are produced by many-body correlations and not single particles excitations (whence my point of different people converging on the same conclusion). The fact that Luttinger liquids cannot have any effective-exciton descriptions might just be a clue that indeed in the solvable, simplest case, the excitonic picture is so bad as even become impossible to draw.
Now assuming that Luttinger liquids don't have any useful notion of exciton to bring to Combescot's formalism—which would be surprising because at the root of it, is still that they go from fermions to bosons and this must include some pairing of some sort—anyway, assuming Luttinger is not good to study this problem, then let us consider this Haldane procedure to treat explicitly an excitonic 1D system, like the one of Lee and Eric Yang[1]. After spending a bit of time with this paper, however, it is my feeling that they are in fact "on our side" of bosonization as understood in our field, and do not rely on anything Haldane-related. I would assume, then, that Combescot would attack them similarly as she does Haug, Hanamura, Schmitt-Rink, Ciuti, you, Ouerdane, Nina, etc. I would also understand that Haldane's procedure cannot be applied in this case.
If the above is correct, then the answer to my initial question is: those are two unrelated techniques bearing the same name. Nothing to discuss.
However I now come to your point 2, which is the one that kept me thinking.
2) Bosonisation in 1D systems in my opinion, is closer to introducing magnons in (ferro/antiferro)magnets or maybe to Frenkel excitons.The magnons in ferro/antiferromagnetic system indeed seem very related to the excitonic conundrum, at least in the way you describe it. And then you say that this was linked to Frenkel excitons (which have also be similarly critically attacked by Combescot and Pogosov[2]) and then you conclude by bringing this back to the Haldane procedure where complicated commutation relations are replaced by much simpler ones.
The answer might be lying somewhere here. I don't see the physical connection—this concept of compositeness, of binding two objects to make a new one, which binding is the root of all problems, this is not present with magnons (you mention two sub-lattices as making it worse but that doesn't create a new object) nor, relatedly, is there any interactions—but I see the mathematical connection, since the starting point of Combescot is the commutator involving another operator $\left[B_m, B_i^\dagger\right] = \delta_{mi} - D_{mi}$. This is why I would feel (I don't know) that Combescot would see no problem with bosonization of such, but if you say that structurally this causes problems for the excited states and that (Haldane's) bosonization applies to them, and that it's in the shape of Combescot's excitons, maybe they are a useful platform to connect Shiva diagrams with bosonization, and find out the connection between them.
Regardingiii) is the connection just being ignored from all sides.
I think that there are people in the field who understand these potential connection much better than we do (at least, much better than I do).
I don't know who would that be. As I told you, top experts from Haldane bosonization had not even heard about the excitonic case.
Anyway, my apologies for a late reply but I've tried to get deeper into the problem and only found myself being quickly out of my depth. The only way to formulate a precise, operation question that I see is to identify one platform that is welcoming of both bosonization à la Haldane, and having excitons, and comparing the outcome of three treatments (along with effective pure-boson Hamiltonian) to see to which extent those do agree or overlap, or indeed contradict each other. This might be Frenkel excitons [does bosonization apply to them?] or magnons [does Combescot's technique applies to their algebra?]. It is also my understanding that you see no important contradiction between Combescot and the rest of the literature, only that she claims them. I would of course be happy to discuss it over zoom, as you suggest.
Probably an important concept in this story is that of self-averaging, as introduced by Lifschitz. It also shows that our problem is related to quenched disorder! Interesting. The all-encompassing, non-additive observables like maximum gap show the worst self-averaging, so are definitely worthy of study!
A more accurate picture of my jaguars:
The curves are the exact $p_\mathcal{C}$, thanks to closed-form shortcut to the exact enumeration which is, combinatorially, out of reach. Here we have $N=1500$. We should look at smaller $N$ too to see when does the smoothness of those curves dies. The points are for smaller sub-sampling: $n=3\times 10^5$. This is actually closer to what I expected. I didn't think one would see purity when embracing the full object, but was foreseeing fluctuations. Obviously! But no. I'll look more into how sub-sampling and $n$ changes that picture. Note that it's only for the donut that we have huge fluctuations anyway. They are quickly negligible, even for small $k$, in general.
As for non-additive observables, we haven't given them any consideration yet, or only anecdotally. These could be the maximal or minimal gap, or in fact the full family $\Delta_{(i)}$ of the $i$th biggest gap: $\Delta_{(1)}=\Delta_\mathrm{max}$, $\Delta_{(2)}$ is the next largest gap, down to $\Delta_{(k)}=\Delta_\mathrm{min}$. It should be interesting to consider $$\Delta_{\left(1\right)}/\Delta_{\left(2\right)}$$ which is dimensionless and scale-free, so should be particularly convenient to sweep over $k$.
Those have no exact subset formula in terms of $P_m$, since $\Delta_\mathrm{max}=\mathrm{max}_c\Delta_c\mathbb{1}_c$ is not a sum, $\mathbb{E}\left[\max\right]\neq\max\mathbb{E}$ and $\Pr\left(\Delta_{\max}<x\right)=\Pr\left(\text{no chord with }\Delta_c\ge x\text{ is an edge}\right) $ is a statement about all the chords jointly. Knowing each $\Pr\left(c\right)$ separately doesn't determine it.
I computed the two-chord level for the variance, and it required a genuinely two-index object $\Pr\left(c\wedge c'\right)$, while skewness would require three-index object $\Pr\left(c\wedge c'\wedge c''\right)$, etc., but the maximum needs it to all orders.
So they are probably the most interesting yet also the least tractable quantities.
I have focused on perimeters, sometimes on areas. But there are plenty of observables of a different nature to consider. The simplest are additive ones,
e.g., additive over points: $$o=\sum_j h\left(\theta_j\right)$$ ($o$ is the observable). Collapse averages then simply read $\mathbb{E}\left[\cdot\mid\mathcal{C}\right]=\left(k/N\right)\sum_{j=1}^{N}h$, with no combinatorics. The most important such observable is maybe the Fourier mode: $$R_{2\ell}=\frac{1}{k}\left\lvert\sum_{j}e^{2i\ell\theta_j}\right\rvert$$ whose mean is set directly by $\sqrt{C_\alpha}$. This serves as an estimator of the state
Then there is then the big family of observables additive over gaps—the ones we have focused on. Calling this time $\delta$ the observable: $$\delta=\sum_i\phi\left(\Delta_i\right)$$
observable $\phi$ note perimeter $\Delta-2\sin\left(\Delta/2\right)$ our focus so far area $\left(\Delta-\sin\Delta\right)/2$ our alternative to perimeters so far Greenwood statistic $\Delta^2$ classic uniformity test, exactly solvable Rényi/Shannon gap entropy $-\Delta\ln\Delta$ sensitive to depletion gaps above a threshold $\Theta\left(\Delta-a\right)$ counts voids directly moments $\sum\Delta^q$ $\Delta^q$ a one-parameter family; $q\to\infty$ interpolates to the maximal gap These should be investigated in turn.
They have the property that the expectations only need the marginal probability of each chord, which we have previously computed as $P_m=\binom{N-m-1}{k-2}\Big/\binom{N}{k}$, so that for an observable of the type $\sum_{\text{edges}}w\left(\text{edge}\right)$, we have $$\mathbb{E}\left[\sum_c w_c\mathbb{1}_c\right]=\sum_c w_c\Pr\left(c\right)=\sum_{i,m}w_{i,m}P_m$$ where $w_{i,m}$ is the chord from the $i$th to the $m$th point of the collapse $\mathcal{C}=\{\theta_i\}$. This is beautiful because in the mean we dissolve the actual sub-sampling of the collapse, it targets the full structure directly, the object itself.
Now for the main statement regarding quantum-state hypothesis. I show here results from single collapses:
The most obvious feature: each collapse is an object even more surely than anything else ever was, because its characterization is smooth—it embeds its own statistics—although it's one single thing, that you can hold in your hands. No ensemble, just the thing, the mountain. And like the jaguars of Gell-Mann, each one has its little characteristics: its smooth, slowly varying departure from the laws of the gender (where are the spots of an individual on its otherwise carpeted fur) imprints a systematic bias that persists across $k$. For each individual, $k$ could be time. It's the order of the observable instead. A collapse that is short at $k=10$ tends to stay short at $k=100$, because the same internal structure (a wide gap somewhere) carries from $k$ to $k$. They can change sign, however, the sign of the collapse's bias is not fixed across scales even though the curve is smooth. You can have a trait in youth that reverts in old age, although it's still the same "you".
This compares the standard quantum averages on the left with the collapse averages on the right.
The right clearly brings some advantages. It is tighter, more symmetric around the mean (the line) and might actually be simpler to compute, or not much more difficult. It merely uses more of the available information. Let me think.
I want to make a statement on Adam and Eve's laws but before that, let's fix notations for the various distributions involved.
We call $D(p|\mathcal{C})$ the distribution of the variate (here the perimeter $p$) for a single, given collapse $\mathcal{C}=\{\theta\}$. By construct, that's for a given $C_\alpha$. This is the unique, mountain-like object, which is reproducible only for those people having access to this collapse.
We call $D_{\mathrm{ens}|\alpha}(p)\equiv\mathbb{E}\big(D(p|C_\alpha)\big)$ the ensemble average over independent collapses but restricted to a given $C_\alpha$.
We call $D_\alpha(p|\mathcal{C})\equiv\mathbb{E}_{\mathcal{C}|\alpha}(D(p|\mathcal{C}))$ the average over distributions of independent collapses for a given $C_\alpha$. It is in red in the plot above.
We call $D_\mathcal{C}(p_\mathcal{C})$ the distribution of the mean $p_\mathcal{C}$ over the various collapses.
We call $D_{\mathrm{ens}}(p)$ the traditional, ensemble average over any state, here for the thermal distribution.
We call $D_\mathcal{C}(p_\mathcal{C}|C_\alpha)$ the distribution of the means obtained from the various collapses of a fixed $C_\alpha$.
We call $D_\mathcal{C}(p_\mathcal{C})$ the distribution of the means obtained from the various collapses over the full ensemble (all $C_\alpha$).
At the level of distributions:
\begin{equation}\label{eq:02wfL}D_{\alpha}(p)\equiv\mathbb{E}_{\mathcal{C}|\alpha}\big(D(p|\mathcal{C})\big)\,,\end{equation} \begin{equation}\label{eq:02YZi}D_{\mathrm{ens}\mid\alpha}(p)=D_{\alpha}(p)\,,\end{equation} \begin{equation}\label{eq:02JTf}D_\mathrm{ens}(p)=\mathbb{E}_\alpha\big(D_{\mathrm{ens}|\alpha}(p)\big)=\mathbb{E}_\alpha\big( D_\alpha(p)\big)=\mathbb{E}_\alpha\big(\mathbb{E}_{\mathcal{C}|\alpha}(D(p|\mathcal{C}))\big)\,.\end{equation} \begin{equation}\label{eq:02wEF}D_{\mathcal{C}}\left(p_{\mathcal{C}}\right)=\mathbb{E}_\alpha\left(D_\mathcal{C}(p_{\mathcal{C}}|C_\alpha)\right)\end{equation} \begin{equation}\label{eq:02NJk}D_\mathcal{C}(p_\mathcal{C}|C_\alpha)\neq D_\alpha(p)\end{equation}
Eq. \eqref{eq:02wfL} is a definition: $D_\alpha$ is the average of the single-collapse laws over independent collapses for the same $C_\alpha$.
Eq. \eqref{eq:02YZi} is a theorem, and follows from a uniformly random $k$-subset of $N$ i.i.d. points being itself $k$ i.i.d. points. It states that two quite different procedures give the same law: one may run independent experiments, each drawing $k$ fresh photons from the same $f_{C_\alpha}$, or collapse $N$ photons, accumulate the narrow mountain-like law from that one collapse, and average over such structures. The two routes differ shot by shot but recover the same final result. The frozen intermediate collapse therefore neither adds nor removes anything: $D_\alpha$ carries no dependence on $N$ at all.
Averaging the collapse laws must not be confused with collecting their means. The latter defines a distribution of a different variate, $D_{\mathcal{C}}(p_\mathcal{C})$, which is narrower than $D_\alpha$.
Eq. \eqref{eq:02JTf} shows that $D_\mathrm{ens}$ is recovered by averaging the previous distributions over the thermal measure of the state sampling.
Eq. \eqref{eq:02wEF} links the homogeneous ensemble to the full ensemble. For the thermal measure, this is $D_{\mathcal{C}}\left(p_{\mathcal{C}}\right)=\int_0^{1/4}\frac{2}{\sqrt{1-4C_\alpha}}\,D_{\mathcal{C}}\left(p_{\mathcal{C}}\mid C_\alpha\,dC_\alpha\right)$.
Eq. \eqref{eq:02NJk} is an inequality and states that the distribution of the variate for a given quantum state $C_\alpha$ differs from the distribution of its means over collapses.
Note that $D(p|\mathcal{C})$ is the only one that depends on $N$. It is, arguably, the most physical object, corresponding to one experiment. Anything else becomes a mathematical description of such raw objects.
So there are various distributions, many objects, and some agree while other do not. We now turn to their mean and other observables and see that agreement increases considerably for the mean but will resist for, say, variances.
The variance can be similarly obtained: $$\mathrm{Var}\left(p_\mathcal{C}\right)=\sum_{i,m}\ell_{i,m}^{2}P_m+2\!\!\sum_{m,m'}\!\!A_s\sum_i \ell_{i,m}\,\ell_{i+m,m'}+\sum_{m,m'}\!\!B_s\!\!\sum_{\substack{i,j\\ \text{disjoint}}}\!\!\ell_{i,m}\,\ell_{j,m'}\;-\;\langle p_\mathcal{C}\rangle^{2}$$ where $\ell_{i,m}\equiv2\sin\left(G_{i,m}/2\right)$ and the probability for the pair of chords $c$ and $c'$ to be realized is $$\Pr\left(c \wedge c'\right)=\begin{cases} {P_m\equiv\binom{N-m-1}{k-2}\big/\binom{N}{k}} & \text{share two endpoints}\\[2pt] A_s\equiv\binom{N-s-1}{k-3}\big/\binom{N}{k} & \text{share one endpoint}\\[2pt] B_s\equiv\binom{N-s-2}{k-4}\big/\binom{N}{k} & \text{fully disjoint}\\[2pt] 0 & \text{overlapping or nested}\end{cases}$$ with $s\equiv m+m'$.
The reasoning is very similar: two chords $c=(i,m)$ and $c'=(j,m')$ can only exist in the $k$-gon if their arcs have disjoint interiors. That kills most pairs and leaves the three cases above, whose probabilities are obtained by the same counting as before: fix the configuration of $c\wedge c'$ and distribute the rest.
Adjacent chords fix $m+m'+1$ points, 3 of which for the two chords. Disjoint chords fix $m+m'+2$ points, 4 of which for the chords. The number of points left to complete the $k$-gon is easily figured out and makes up the combinatorics of what is available. Same reasoning, really. The final formula gets assembled by weighting the contributing terms to $\langle p^2\rangle_\alpha$.
Some results regarding the mean of single-collapse observables. In particular, the expression from the mean perimeter for a given $C_\alpha$, directly from the points of a single collapse as opposed to deriving it from the distribution, is surprisingly simple: $$p_\mathcal{C}=2\pi-\sum_{i=1}^{N}\sum_{m=1}^{N-k+1}{\binom{N-m-1}{k-2}\over\binom{N}{k}}\left[G_{i,m}-2\sin\left(G_{i,m}/2\right)\right]$$ where $G_{i,m}\equiv\Delta_{i+1} + \Delta_{i+2} + \cdots + \Delta_{i+m}$ the $i$th sum of the $m$ consecutive arcs (distances on the circle) of the points. This is because the averaging dissolves much of the structure: a single realization must know which chords occur. The perimeter becomes a sum of strongly correlated indicator-weighted terms, and to get $D(p)$, one needs the joint law of all of them. That is the hard object. The average, on the other hand, is oblivious to this joint structure, as it flies through the sum to kill the overly specific details into a probability: $$\mathbb{E}\left[\sum_c \ell_c \mathbb{1}_c\right] = \sum_c \ell_c \Pr(c)$$ where $\ell_c$ is the length of the chord and $\mathbb{1}_c$ the indicator function: 1 if the chord is in the $k$-gon, 0 otherwise. So we don't need to do any counting of exactly $k$ points: on average, that'll come out the same. We only need to compute the probability of a given chord.
The overall result is obtained from the "mean deficit", which is what the polygon takes away from the circle. The chord itself $\ell_c$ is, for the $i$th point connected to the $(i+m)$th one, $G_{i,m}-2\sin\left(G_{i,m}/2\right)$ since the arc $G$ gives a chord $2\sin\left(G/2\right)$, so the deficit is their difference.
The probability $\Pr(c)$ is more challenging: it is obtained by fixing the chord and counting how many $k$-gons are left. A particular chord removes $m+1$ points from the pool: the starting one, $m-1$ points in between that are skipped over, and the ending point of the chord. This takes 2 points from the $k$-gon. From the $N-(m+1)$ points remaining, one has $k-2$ points to distribute to complete the $k$-gon. So for a fixed chord that took $m+1$ points, the probability that it will be realized, that is, the probability that it belongs to a $k$-gon, is how many such $k$-gons [with this chord] can be formed over how many $k$-gons can be formed at all. The former is ${N-(m+1)\choose k-2}$. The latter is ${N\choose k}$. Their ratio is the probability $\Pr(c)$. That completes the expression above.
Analysis of the single-collapse heterogeneity for the atomic case $C_\alpha=1/6$. This one is emphatically not fractal. Is it because it is the atom? Probably not, but that the atom is not fractal might still be relevant.
As I'm writing the text («This tells us that each collapse, for some high multiphoton observables, become a reality of its own, a particular instances of a scenario which, in its own complexity, can only be described on a per-individual basis, empirically»), I'm looking simultaneously at the quantum-state ergodic hypothesis of high-multiphoton states.
This is the distribution of dihectapentacontagon (250-gons) out of 500 photon thermal bimodal vortices:
What is important here is that each distribution is markedly different from the other. But we're not looking at noise: each distribution is the same regardless of which sampling of points one will make from each collapse to derive the distribution. And there's a lot of choice: $\binom{500}{250}\approx10^{149}$, so one has to sub-sample from that. Whatever sub-sampling you do, you get the same distribution. This is shown here:
So each ugly duckling-distribution is a real object of its own. Like a mountain, each is different, each is unique, each is special.
Now the beautiful part, while the distributions are different, their most important statistical quantifier—the mean—is not. It fluctuates according to this skewed distribution, but on average, it is the result of quantum mechanics. We're speaking about the mean of the means, here:
This is an exact result that is independent of $C_\alpha$. Here is the case for the most correlated (classical) case $C_{1/4}$. Another striking thing for this case is that the distributions can really become accidented, fractal looking, yet still reproducible, tangible objects. I show you three different sampling, of three collapses now to make extra sure of this:
We could give name to those things—like satellites (Kerberos, Styx and Nix), or AF tags: 02YsC (af), 02lFL (af) and 02wWT (af). They are truly unique objects.
As indeed, the distributions of means show the same result: the mean of the means is the theoretical one, despite the fractality, despite the completely different distribution:
It's beautiful. Who would think such fractally-accidented profiles give you the mean you want? (in the mean); It's meta-statistics that we're doing here, and I love it. The more into particular cases you go, the higher in the ladder of abstraction you need to climb to make it general again!
The distribution of perimeters for a four-photon quadrilateral (projected on the unit circle) is of the type $D_{C_{\boldsymbol\alpha}}(p)=A_0(p)+C_{\boldsymbol\alpha}\,A_1(p)+C_{\boldsymbol\alpha}^2\,A_2(p)$ where $A_0$ is the distribution without correlations (the donut) and $A_i$, $i=1,2$, are fluctuations or deviations around this to account for "correlations" for the given $C_\boldsymbol{\alpha}$. The moments calculated from this distribution will thus invoke $\langle C_\boldsymbol{\alpha}\rangle$ and $\langle C_\boldsymbol{\alpha}^2\rangle$ and since $\langle C_\boldsymbol{\alpha}^2\rangle\neq\langle C_\boldsymbol{\alpha}\rangle^2$, we need another atom to provide for the added correlation.
What is surprising is the quality of the one-atom approximation. You really need a lot of photons $N\gg1$ to see by eye the one-atom approximation fail. Here's an overview for a few representative cases:
What are we looking at here?
- in (a), the red is exact (matches the black which is over the full thermal ensemble) but the red is off because it misses one atom. However, to the eye, it looks perfect.
- in (b) we see the failure of the green.
- in (c) we see the agreement of the red. It doesn't look much better than the actually failing green.
- in (d) we have (c) again in a different color as this column shows the case with enough atoms, which is matched on the first row.
- in (e) the case of $D_8$ requiring three atoms. This time the green disagreement is a bit visible by eye.
- in (f) the one-atom failure for 8-photon observable is clear.
- in (g) the two-atom failure for 8-photon observable is not clear (expected, purple). The red seems to be as good as (h).
- in (h) the three-atom working case for 8-photon observables. Not clearly better than the 2-atom one.
- in (i) the $D_{40}$ failure of one-atom is obvious, the two-atoms remains excellent
- in (m) the $D_{100}$ confirms the pattern.
We can see that a few atoms allow a very good agreement with even very-high multiphoton observables.
And this one comes from Jacob. What we see here is the distribution of means, each computed from a one-collapse distribution of distances; each collapse had 200 points (he told me) and that's how they compare to the theoretical means for (left) donut, (center) atom = thermal and (right) dipole. As expected, they agree. Now we should see that for a very-high multiphoton observable, where distributions will become fractal, and unique. But their mean could (should) still agree.
A first surprise is that the two-atom necessity is far from obvious for four-photon observables (the first to require two atoms). That's what I find, and initially thought there was a mistake in my atomicity requirements, but no, as I discuss later. First I want to see if such a very good agreement remains for higher observables (at some point it should break noticeably).
Moment of weakness of the day as I see I've been loosing time looking into this myself instead of asking Claude to solve everything from the start. The basic idea is that the atoms are actually Gauss-Legendre nodes once your regularise the thermal distribution into a flat one, and that allows you to reach high numbers, with which we could produce this following progression of how the thermal ensemble is being populated "canonically" [always choosing the most-covering moments distributions].
There's a more crowded uniform-looking coverage of the full ensemble, striving to touch the extremes (in particular the most-correlated dipole) but failing to. Nothing unexpected or very surprising but a beautiful snapshot of quantum vs ensemble averages.
The derivation goes as follows: the thermal correlator is $C_{\boldsymbol\alpha}=u(1-u)$ with $u$ uniform on $[0,1]$ so that corresponding measure is $$\mu(C_{\boldsymbol\alpha})=\frac{2}{\sqrt{1-4C_{\boldsymbol\alpha}}}\ ,\qquad C_{\boldsymbol\alpha}\in\left[0,\tfrac14\right]\,.$$ Now Claude observes that regularizing that, $$t\equiv1-2u\ ,\qquad C_{\boldsymbol\alpha}=\frac{1-t^{2}}{4}\ ,\qquad t\in[-1,1]$$ then the measure becomes flat: $t$ is uniform with density $\tfrac12$. Then, it claims, the atoms are the Gauss quadrature of a constant weight on $[-1,1]$, which brings it to our polynomial but now in Legendre form: $$\pi_\nu\!\left(C_{\boldsymbol\alpha}\right)\ \propto\ P_{2\nu}\!\left(\sqrt{1-4C_{\boldsymbol\alpha}}\right)$$ where $P_{2\nu}$ is the ordinary Legendre polynomials on $\left[-1,1\right]$, $P_0=1$, $P_1=t$, $P_2=\tfrac12\left(3t^{2}-1\right)$, and so on.
The atoms follow immediately. The roots of $P_{2\nu}$ come in pairs $\pm s_k$, and both members map to the same location, so the $2\nu$ Legendre nodes collapse onto $\nu$ atoms. Their weights are equal by symmetry, and the density $\tfrac12$ converts the Legendre total of $2$ into unity. Hence, with $s_k$ the positive nodes of the $2\nu$-point Gauss-Legendre rule and $\lambda_k$ their standard weights, we get all our atoms easily: $$\boxed{C_k=\frac{1-s_k^{2}}{4}\ ,\qquad w_k=\lambda_k\,.}$$
Nothing is left to solve. No Hankel matrix, no root-finding. The first three cases:
$\nu$ serves $N$ $s_k$ $C_k$ $w_k$ 1 1–3 0.5773502692 0.1666666667 1.0000000000 2 4–7 0.3399810436 0.2211032225 0.6521451549 0.8611363116 0.0646110632 0.3478548451 3 8–11 0.2386191861 0.2357652210 0.4679139346 0.6612093865 0.1407005368 0.3607615730 0.9324695142 0.0326251513 0.1713244924 4 12–15 0.1834346425 0.2415879330 0.3626837834 0.5255324099 0.1809539215 0.3137066459 0.7966664774 0.0913306309 0.2223810345 0.9602898565 0.0194608479 0.1012285363 5 16–19 0.1488743390 0.2444591078 0.2955242247 0.4333953941 0.2030421081 0.2692667193 0.6794095683 0.1346006596 0.2190863625 0.8650633667 0.0629163429 0.1494513492 0.9739065285 0.0128765184 0.0666713443
The atoms are basically moment-generating machines. A given number of them can produce more than are needed by a given observable. I got something badly wrong earlier, not appreciating that fully, which is how many atoms we need to cover $N$-photon observables. I thought you need three for $N=6$, but no, you're still good with two atoms. Only, you cannot choose them. You don't have enough moment-freedom to move them around. Let us call "slack" (for margin, or looseness in a rope) the excedent of moments we get.
$$\text{slack}=\left(2\nu-1\right)-M$$
which is either 0 or 1. Whenever it is 0 (no slack), the atoms are uniquely defined. When it is 1, the extra freedom allows to settle for other atoms. So this was the case in 00W0b, there I was giving a family of solutions for $N=4$ and $5$. In the previous post, applying strictly the procedure, which is the one that cuts the slack, I also provide the soutions for $N=6$ and $7$. Those are also valid for $N=4$ and $5$, which, however, have more pairs of atoms.
So the photon-observables served by the same number of atoms go by groups of 4:
- 0-3 photon observables with 1 atom (a 0 observable is no observable but that's the pattern)
- 4-7 photon observables with 2 atoms
- 8-11 photon observables with 3 atoms
etc.
The polynomial for 3 atoms by the way is $$924x^{3}-378x^{2}+42x-1$$ with solutions $$C_k=\frac{3}{22}+\frac{\sqrt{15}}{33}\,\cos\left(\frac{1}{3}\arccos\left(-\frac{\sqrt{15}}{35}\right)-\frac{2\pi k}{3}\right)\ ,\qquad k=0,1,2$$ with corresponding weights $$w_k=\frac{660\,C_k^{2}-160\,C_k+7}{30\left(66\,C_k^{2}-18\,C_k+1\right)}\,.$$
Generalization to higher atoms would maybe look futile, except that Claude found another way that seems to be much more efficient than all the previous material.
Now back to particular cases, this time constructive. The procedure as I sketched it has three stages:
- get $\nu$ from $N$.
- solve the Hankel equations for $\pi_\nu$ and take its roots as the locations.
- solve the Vandermonde equations for the weights.
I now do this for $N\le7$, which improves by two additional photons my previous solution and highight and important point to which I come back in the next post, as it had escaped me.
We remind the moments $m_p=\int_0^1u^{\,p}\left(1-u\right)^{p}du=\frac{p!^{2}}{\left(2p+1\right)!}$ so $\left\{\frac{1}{6},\frac{1}{30},\frac{1}{140},\frac{1}{630},\frac{1}{2772}, ...\right\}$
Cases $N=2$ and $N=3$: In this case $M=\lfloor N/2\rfloor=1$ and $\nu=\left\lceil\left(M+1\right)/2\right\rceil=1$. The Hankel equation for $r=0$ only yields $m_0\,a_0=-m_1\ \Longrightarrow\ a_0=-\frac{m_1}{m_0}=-\frac16$. The polynomial is thus of order 1 $\pi_1\left(x\right)=x+a_0=x-\frac16$ with root $C_1=1/6$ which is the value found by Daniel. We don't have to solve the Vandermonde has the weights has to one.
Cases $N\in\{4, 5\}$: $M=2$ and $\nu=\left\lceil3/2\right\rceil=2$ atoms. But this is also for the Cases $N\in\{6,7\}$ in which case $M=3$ but $\nu=\left\lceil4/2\right\rceil$ is still two atoms. This difference is discussed in the next post. The following applies to all four cases.
The Hankel equation goes with $r=0$ and $r=1$ as: $$\begin{pmatrix} m_0 & m_1\\ m_1 & m_2\end{pmatrix}\begin{pmatrix} a_0\\ a_1\end{pmatrix}=-\begin{pmatrix} m_2\\ m_3\end{pmatrix}\ \Longrightarrow\ \begin{pmatrix} 1 & 1/6\\ 1/6 & 1/30\end{pmatrix}\begin{pmatrix} a_0\\ a_1\end{pmatrix}=-\begin{pmatrix} 1/30\\ 1/140\end{pmatrix}$$ which gives the polynomial as $$\pi_2\left(x\right)=x^{2}-\frac27x+\frac{1}{70}\ \Longleftrightarrow\ 70x^{2}-20x+1=0$$ with two locations: $$C_{\pm}=\frac{10\pm\sqrt{30}}{70}$$ i.e., $C_+\simeq0.22$ and $C_-\simeq0.06$. Both lie inside $\left[0,\tfrac14\right].$ The Vandermonde equation reads: $$\begin{pmatrix} 1 & 1\\ C_+ & C_-\end{pmatrix}\begin{pmatrix} w_+\\ w_-\end{pmatrix}=\begin{pmatrix} 1\\ 1/6\end{pmatrix}$$
$C_j$ $w_j$ $j=+$ $\left(10+\sqrt{30}\right)/70$ $\left(18+\sqrt{30}\right)/36$ $j=-$ $\left(10-\sqrt{30}\right)/70$ $\left(18-\sqrt{30}\right)/36$ This is a particular case and the question of which is the general is more interesting and subtle than I thought.
Now for the general case, although still for the phase-invariant case, I provide the equations to solve to get the weights and locations of the atoms.
We had left it at these $M+1$ equations with $2\nu$ unknowns: \begin{equation}\label{eq:00uIJ}\sum_{j=1}^{\nu}w_j\,C_j^{\,p}=m_p\end{equation} for $m_p\equiv\langle C_{\boldsymbol\alpha}^{\,p}\rangle$ and $p=0,1,\dots,M$. Since $M$ is the highest power of $C_{\boldsymbol\alpha}$ needed to describe the state, and in the case of phase-invariance, the unbalanced terms are gone, leaving us with only $p=q$, and since $p+q$ cannot exceed $N$ this caps $p$ at $M=\lfloor N/2\rfloor$. As for $\nu$, this is the number of atoms minimally needed, each coming with two parameters: the weight (mass in the traditional lexicon) and the location (the quantum state for us). We're looking for the $\nu$ pairs $(w_j,C_j)$ for $1\le j\le\nu$. We determine its size below.
We need to solve for $w_j$ the weights and $C_j$ the quantum-state coefficients, or "locations" (of the atoms), whose $p$th powers are weighted by the weight to produce the averages. Each $(w_j,C_j)$ provides an atom for the distribution. There are $M=\lfloor N/2\rfloor.$
These are linear equations in the weights. We solve first for the locations. The proof is constructive, so if we have a solution, it will be enough. Let us thus assume that the locations are algebraic numbers, i.e., that they are solution of a polynomial $$\pi_\nu(x)=x^{\nu}+a_{\nu-1}x^{\nu-1}+\cdots+a_0\ ,\qquad \pi_\nu(C_j)=0\ \text{for all}\ j\,.$$ Multiplying Eq. \eqref{eq:00uIJ} at index $p=s+r$, i.e., $\sum_{j=1}^{\nu}w_j\,C_j^{\,s+r}=m_{s+r}$, by $a_s$ and summing over $s$ from $0$ to $\nu$, $$\sum_{j=1}^{\nu}w_j\sum_{s=0}^{\nu}a_s\,C_j^{\,s+r}=\sum_{s=0}^{\nu}a_s\,m_{s+r}$$ but $\sum_{s=0}^{\nu}a_s\,C_j^{\,s+r}=C_j^{\,r}\sum_{s=0}^{\nu}a_s\,C_j^{\,s}=C_j^{\,r}\,\pi_\nu(C_j)=0$, and therefore the right-hand side vanishes too. Isolating the highest-order term, which is $s=\nu$ with $a_\nu=1$: $$m_{\nu+r}+\sum_{s=0}^{\nu-1}a_s\,m_{s+r}=0\ ,\qquad r=0,1,\dots,\nu-1\,.$$ This is a linear system for the coefficients $a_s$ involving only the moments $m$, and it is in the Hankel form since the entry at position $(r,s)$ is $m_{s+r}$. In matrix form:
$$\mathbf{M}_\nu\,\mathbf{a}=-\mathbf{h}\,,$$ where $\left(\mathbf{M}_\nu\right)_{rs}=m_{r+s},$ with $ r,s=0,\dots,\nu-1,$ $\mathbf{a}\equiv\left(a_0,a_1,\dots,a_{\nu-1}\right)^{\mathsf T}$ and $\mathbf{h}=\left(m_\nu,m_{\nu+1},\dots,m_{2\nu-1}\right)^{\mathsf T}$. Explicitly: $$\begin{pmatrix} m_0 & m_1 & \cdots & m_{\nu-1}\\ m_1 & m_2 & \cdots & m_{\nu}\\ \vdots & \vdots & & \vdots\\ m_{\nu-1} & m_{\nu} & \cdots & m_{2\nu-2}\end{pmatrix}\begin{pmatrix} a_0\\ a_1\\ \vdots\\ a_{\nu-1}\end{pmatrix}=-\begin{pmatrix} m_{\nu}\\ m_{\nu+1}\\ \vdots\\ m_{2\nu-1}\end{pmatrix}\,.$$ And that's where we find the magnitude of $\nu$. The largest value to enter is $m_{2\nu-1}$ in the last entry of the right-hand side $\mathbf{h}$ (and not in $\mathbf{M}_\nu$ where this is $m_{2(\nu-1)}$). We start from $m_0$ and need all the moments in between those two, so $2\nu$ of them, corresponding to the number of unknowns for the atoms. The top indices should match but $2\nu-1$ is odd while $M$ can be anything, therefore we have to settle with the smallest $\nu$ such that $2\nu-1\ge M$, which has solution $\nu=\left\lceil\frac{M+1}{2}\right\rceil$. Substituting $M=\lfloor N/2\rfloor$ gives $\nu=\left\lceil\left(\lfloor N/2\rfloor+1\right)/2\right\rceil$, and the nested floor-ceiling collapses to $$\nu=\left\lceil{N+1\over4}\right\rceil\,.$$
As a Hankel matrix, its antidiagonals are constant, with $\det \mathbf{M}_\nu>0$ so invertible, with also real and distinct roots. Their belonging to the support is unclear to me mathematically but we expect it is the case for semi-classical states and not so for quantum ones. So we solve for two things: first for the coefficients $a_s$ of the polynomial equations, and then for the roots of those equations which provide the locations of the atoms.
Then we have to solve for the weights, but those are from a linear set of equations, so this is, in principle, a formality. This comes out as a Vandermonde system: $$\begin{pmatrix} 1 & 1 & \cdots & 1\\ C_1 & C_2 & \cdots & C_\nu\\ \vdots & \vdots & & \vdots\\ C_1^{\nu-1} & C_2^{\nu-1} & \cdots & C_\nu^{\nu-1}\end{pmatrix}\begin{pmatrix} w_1\\ w_2\\ \vdots\\ w_\nu\end{pmatrix}=\begin{pmatrix} m_0\\ m_1\\ \vdots\\ m_{\nu-1}\end{pmatrix}$$ which is nonsingular if the $C_j$ are non-degenerate, which they should.
A beautiful construction: locations from Hankel, weights from Vandermonde, both guaranteed physical, hence the representation exists with $\left\lceil\left(N+1\right)/4\right\rceil$ atoms.
Before I work the general case, the particular cases: for $N=2$ and $N=3$, $C_j=1/6$ indeed (as found by Daniel yesterday) and $w_1=1$ (one atom has all the weight).
The general moment is $m_p=(p!)^2/(2p+1)!$, i.e. $1,\ 1/6,\ 1/30,\ 1/140,\dots$, so for $4\le N\le 5$ (two next cases), then, the possible various minimal solutions are of the type: (minimal because one can take more atoms, mind you, you can still take the full thermal bag): $$C_2=m_1-\frac{\sigma^{2}}{C_1-m_1}\ ,\qquad w_1=\frac{\sigma^{2}}{\sigma^{2}+\left(C_1-m_1\right)^{2}}\ ,\qquad w_2=1-w_1$$
$C1$ $w_1$ $C_2$ $w_2$ $1/5$ $5/6$ $0$ $1/6$ $\left(10+\sqrt{30}\right)/70$ $0.652145$ $\left(10-\sqrt{30}\right)/70$ $0.347855$ $1/4$ $4/9$ $1/10$ $5/9$ where the moment is $1/30$ and $\sigma^2\equiv\mathrm{Var}\left(C_{\boldsymbol\alpha}\right)=m_2-m_1^2=\frac{1}{30}-\frac{1}{36}=\frac{1}{180}\,.$ It's because one atom alone cannot account for the four or five particle fluctuations that you need another one.
One of the extrema as $C_1=1/4$, i.e., the dipole (most classically correlated). But this is contaminated by the $C_2=1/10$, which also has higher probability.
In the heterogeneous samples boson-spatial-correlations problem, the one-distribution-does-it-all result is a problem of $\nu$-atom replacements ("atom" in the sense of measure theory). This will bring us to express \begin{equation}\label{eq:00zPI}\overline{\tilde\rho^{(N)}}(\{\theta\})=\sum_{j=1}^{\left\lceil\frac{N+1}{4}\right\rceil}w_j\int_0^{2\pi}\frac{d\phi}{2\pi}\prod_{i=1}^{N}\tilde\rho^{(1)}_{\sqrt{C_j}e^{i\phi}}(\theta_i)\end{equation} where the $\nu\equiv\left\lceil\frac{N+1}{4}\right\rceil$ atoms provide the location $C_j$ and mass $w_j$ as we will show in the following and following post.
From: $$\tilde\rho^{(N)}_{\boldsymbol\alpha}(\theta_1,\dots,\theta_N)=\prod_{i=1}^N\tilde\rho^{(1)}_{\boldsymbol\alpha}(\theta_i)$$
with the one-photon distribution of LG$\pm\ell$ $$\tilde\rho^{(1)}_{\boldsymbol\alpha}(\theta)=\frac{1}{2\pi}\left[1+\gamma e^{2i\ell\theta}+\gamma^*e^{-2i\ell\theta}\right]$$ where $$\gamma\equiv\frac{\alpha_a\alpha_b^*}{\lvert\alpha_a\rvert^2+\lvert\alpha_b\rvert^2}\ ,\qquad \lvert\gamma\rvert^2=C_{\boldsymbol\alpha}\,.$$ The four real parameters of $\boldsymbol\alpha$ enter only through the single complex number $\gamma$, so from here on we index the densities by $\gamma$ and write $\tilde\rho^{(1)}_{\gamma}$ for $\tilde\rho^{(1)}_{\boldsymbol\alpha}$. The structure of this becomes, introducing $z_i\equiv e^{2i\ell\theta_i}$ \begin{equation}\label{eq:00zxQ} (2\pi)^N\,\tilde\rho^{(N)}_{\boldsymbol\alpha}(\theta_1,\dots,\theta_N)=\prod_{i=1}^N\left(1+\gamma z_i+\gamma^*z_i^*\right)=\sum_{p,q\atop p+q\le N}\gamma^p\gamma^{*q}\,S_{pq}(\{\theta\}) \end{equation} where $S_{pq}$ is our long-term combinatoric friend $$S_{pq}(\theta_1,\dots,\theta_N)\equiv\sum_{\substack{P,Q\subseteq\{1,\dots,N\},\ P\cap Q=\varnothing\\ \lvert P\rvert=p,\ \lvert Q\rvert=q}}\exp\left[2i\ell\left(\sum_{i\in P}\theta_i-\sum_{j\in Q}\theta_j\right)\right]$$ which has $\binom{N}{p}\binom{N-p}{q}$ terms (cf. the Distribution of bosons paper).
So far it's only algebra. We did nothing. Now let's turn to quantum states, this will introduce a paradigm shift of retaining a wavefunction, while having performed an average over the quantum states already. For thermal, for instance, $C_{\boldsymbol\alpha}=u(1-u)$ with $u$ uniform so that the thermal measure reads $$\mu(C_{\boldsymbol\alpha})=\frac{2}{\sqrt{1-4C_{\boldsymbol\alpha}}}\ ,\qquad C_{\boldsymbol\alpha}\in\left[0,\tfrac14\right]\,.$$ This average is $\overline{\dots}\equiv\int d\boldsymbol\alpha\,P(\boldsymbol\alpha)$, in which case
$$(2\pi)^N\,\overline{\tilde\rho^{(N)}}=\sum_{p+q\le N}\langle\gamma^p\gamma^{*q}\rangle\;S_{p,q}(\{\theta\})\,,$$
Let's consider our de Finetti representation:
$$\overline{\tilde\rho^{(N)}}(\{\theta\})=\int \prod_{i=1}^{N}\tilde\rho^{(1)}_{\gamma}(\theta_i)d\mu(\gamma)$$ with $\mu$ the distribution of $\gamma$ over the ensemble defined by the quantum state. We now look for a second measure $\mu'$ which gives the same result in terms of atoms. If it exists (and we come to that in a next post), then by assumption: $$\mu'=\sum_{j=1}^{\nu}w_j\,\delta_{\gamma_j}\quad\text{and }\quad\overline{\tilde\rho^{(N)}}=\sum_{j=1}^{\nu}w_j\prod_{i=1}^{N}\tilde\rho^{(1)}_{\gamma_j}(\theta_i)$$ so we have
\begin{equation}\label{eq:00zDS}(2\pi)^N\sum_{j=1}^{\nu}w_j\prod_{i=1}^{N}\tilde\rho^{(1)}_{\gamma_j}(\theta_i)=\sum_{p+q\le N}\langle\gamma^p\gamma^{*q}\rangle\;S_{p,q}(\{\theta\})\,.\end{equation}
At the same time, applying Eq. \eqref{eq:00zxQ} to the case of the atom $\gamma_j$, which is a fixed complex number like any other, we have $$\left(2\pi\right)^{N}\prod_{i=1}^{N}\tilde\rho^{(1)}_{\gamma_j}(\theta_i)=\sum_{p+q\le N}\gamma_j^{\,p}\gamma_j^{*q}\,S_{pq}(\{\theta\})$$
so that after multiplying by $w_j$, summing over $j$ and equating to Eq. \eqref{eq:00zDS}, we find that we must have (because the $S_{pq}$ are linearly independent, which should also be proven but I'll take it as given): \begin{equation}\label{eq:00ECG}\sum_{j=1}^{\nu}w_j\,\gamma_j^{\,p}\gamma_j^{*q}=\langle\gamma^{p}\gamma^{*q}\rangle\ ,\qquad p+q\le N\,.\end{equation}
Let us assume now the particular case of phase-invariance, in which case the atoms become not points but rings (i.e., phase averaged), each $\delta_{\gamma_j}$ being replaced by the uniform measure on the circle of radius $\sqrt{C_j}$—which is the form that was assumed in Eq. \eqref{eq:00zPI} that should eventually be generalized to non-phase-invariant cases—and Eq. \eqref{eq:00ECG} simplifies to $$\sum_{j=1}^{\nu}w_j\,C_j^{\,p}=m_p\ ,\qquad\text{where } m_p\equiv\langle C_{\boldsymbol\alpha}^{\,p}\rangle\ ,\qquad p=0,1,\dots,M$$ with $M$ the largest $p$ allowed, that is $M=\lfloor N/2\rfloor$. This gives us $M+1$ equations with $2\nu$ unknowns. I address the problem of solving this separately. The simplest one gives us the case we discovered yesterday during the meeting. The big result for now is that this depends on the order of the observable, with a step-wise increase that is rank-2 specific, which is Eq. \eqref{eq:00zPI}, that is not yet fully proven, since in addition to providing an explicit construction for $w_j$ and $\gamma_j$ (finding the atoms), we must also justify the $\left\lceil\frac{N+1}{4}\right\rceil$, which I'll do now.
I must summarize the past days study of uniformity/homogeneity of CSI violating rank-2 fields led me to this perspective between our case and that of older SPDC literature, which was the question that started this long detour: $$\text{uniform }G^{(1)}\ \wedge\ \text{homogeneous }G^{(2)}\ \wedge\ \text{rank }2\ \Longrightarrow\ \text{CSI never violated}$$ I provided pairwise cases (uniform and violating; homogeneous and violating; something uniform & homogeneous is trivial but non-violating as per the theorem).
Now for the 3rd, maybe most ambitious characterization of local violation of CSI: the direction of maximum violation: It is the unit vector which makes the angle $\theta_\mathrm{max}$ (with respect to $\hat{\mathbf{x}}$) where $$\theta_\mathrm{max}={1\over2}\arctan\left({2(g_xg_y-p_xp_y)\over g_x^2-p_x^2-g_y^2+p_y^2}\right)$$ and the corresponding maximum value of this violation is: $$Q_{\max}\equiv\lambda_+(M)=\tfrac12\big[T+\sqrt{T^2+4(g\wedge p)^2}\big]$$ where $\lambda_+$ is the positive-value eigenvalue of $M$ (which is indefinite, so it has one $\ge0$) and $T\equiv|g|^2-|p|^2$. This gives the anisotropy of the landscape.
Proof: To find the direction where $Q(\hat{\boldsymbol\delta})=\hat{\boldsymbol\delta}^{\mathsf T}M\,\hat{\boldsymbol\delta}$ increases the most, we diagonalize its $M=\lambda_+\hat e_+\hat e_+^{\mathsf T}+\lambda_-\hat e_-\hat e_-^{\mathsf T}$, with $\hat e_\pm$ orthonormal. Now, writing $\hat{\boldsymbol\delta}=\cos\psi\,\hat e_++\sin\psi\,\hat e_-$, we have
$$ Q(\hat{\boldsymbol\delta})=\lambda_+\cos^2\psi+\lambda_-\sin^2\psi=\lambda_-+(\lambda_+-\lambda_-)\cos^2\psi $$
which is a convex combination of the two eigenvalues, so $\lambda_-\le Q\le\lambda_+$, with the maximum at $\psi=0$, i.e. $\hat{\boldsymbol\delta}=\hat e_+$. The direction of highest $Q$ (CSI violation) is thus the eigenvector of $\lambda_+$, and $Q_{\max}=\lambda_+$.
So we are left with the eigenproblem of $M$. For a $2\times2$ symmetric matrix the characteristic polynomial is $\lambda^2-T\lambda+D=0$ where $T\equiv\operatorname{tr}M$ and $D\equiv\det M$, so
$$ \lambda_\pm=\frac{T\pm\sqrt{T^2-4D}}{2}\,.$$
We already previously calculated $T=|g|^2-|p|^2$ and $D=-(g\wedge p)^2$, so that $ \lambda_\pm=\frac{1}{2}\left[\big(|g|^2-|p|^2\big)\pm\sqrt{\big(|g|^2-|p|^2\big)^2+4(g\wedge p)^2}\right]$, where $g\wedge p=g_xp_y-g_yp_x$, or explicitly, as the $z$ component of the vector product of the gradients of $\Lambda$ and $\varphi$: $$g\wedge p=\partial_x\Lambda\,\partial_y\varphi-\partial_y\Lambda\,\partial_x\varphi\,.$$
19:44 (CET).
The map of local violation in the previous post counts directions, weighting a barely-violating one the same as a strongly-violating one. We can quantify the amount of violation by average its amount over all angles. That gives us: $$\boxed{\langle Q\rangle=\tfrac12\big(|g|^2-|p|^2\big)=\tfrac12\big(|\nabla\Lambda|^2-|\nabla\varphi|^2\big)}\,.$$
16:44 (CET).
So, now, let me look at ways to see the local violations of arbitrary cases. The first one will be: $$\boxed{f(\mathbf r)=1-\frac{1}{\pi}\arccos\frac{|g|^2-|p|^2}{|g+p|\,|g-p|}\in[0,1]\,.}$$
This quantifies how much around of each point there is a violation. At $\mathbf{r}_1=\mathbf{r}_2$ we always have saturation, of course. But if you step a little bit, then you can, or not, violate. Where? Let's keep introducing new notations until we gag: so now give me $u\equiv g+p$ and $v\equiv g-p$. There is CSI violation iff $$Q=(\hat{\boldsymbol\delta}\cdot u)(\hat{\boldsymbol\delta}\cdot v)>0$$ i.e.,—as already noted—they should have the same sign.
So let's assume $\hat{\boldsymbol\delta}\cdot u>0$. That happens for the full half-circle centered around $u$ (which is a vector, of course, the sum of $g$ and $p$ themselves the gradients of logs of amplitudes and phase yada yada yada). A big of planar geometry now: so we have two vectors with an angle $\Delta$ between them. Two arcs of length $\pi$ whose centres are separated by the angle $\Delta$ between $u$ and $v$ overlap on an arc of length $\pi-\Delta$. The both-negative case is the antipodal copy, and provides another $\pi-\Delta$. So out of the full $2\pi$, we have $f=\frac{2(\pi-\Delta)}{2\pi}=1-\frac{\Delta}{\pi}$ of same-sign quantities and thus CSI violation. This angle $\Delta$ we can get from the law of cosines: $$\cos\Delta=\frac{u\cdot v}{|u||v|}=\frac{(g+p)\cdot(g-p)}{|g+p||g-p|}=\frac{|g|^2-|p|^2}{|g+p||g-p|}\,.$$ Neat, no? So that defines a first useful 2D map of local violations, which is the boxed formula above.
16:31 (CET).
The differential form of $\cosh\Delta\Lambda+\cos\Delta\varphi>2$ reads $1+(\boldsymbol\delta\cdot\nabla\Lambda/2)^2+1-(\boldsymbol\delta\cdot\nabla\varphi/2)^2>2$, i.e., with $\hat{\boldsymbol\delta}\equiv\boldsymbol\delta/|\boldsymbol\delta|$ the unit vector of distance separation:
\begin{equation}\label{eq:00iBv}(\hat{\boldsymbol\delta}\cdot\nabla\Lambda)^2-(\hat{\boldsymbol\delta}\cdot\nabla\varphi)^2>0\,.\end{equation}
Calling $g\equiv\nabla\Lambda$ the direction of steepest increase of the log amplitude ratio $\Lambda(\mathbf{r})\equiv\ln\big|\phi_b(\mathbf{r})/\phi_a(\mathbf{r})\big|$ and $p\equiv\nabla\varphi$ the direction of steepest increase of the relative phase $\theta_b-\theta_a$, the CSI violation criterion Eq. (\ref{eq:00iBv}) becomes $$Q(\hat{\boldsymbol\delta})\equiv\big(\hat{\boldsymbol\delta}\cdot(g+p)\big)\big(\hat{\boldsymbol\delta}\cdot(g-p)\big)$$ which should be $>0$ for violation, i.e., they should have the same sign. But since, in this form, there exists $M$ such that $$Q(\hat{\boldsymbol\delta})=\hat{\boldsymbol\delta}^{\mathsf T}M\,\hat{\boldsymbol\delta}$$ with $$M\equiv gg^{\mathsf T}-pp^{\mathsf T}=\begin{pmatrix}g_x^2-p_x^2 & g_xg_y-p_xp_y\\ g_xg_y-p_xp_y & g_y^2-p_y^2\end{pmatrix}\,.$$ Since $\det M=-(g\wedge p)^2\le0$ (where $\wedge$ is the $z$-component of the vector product), it is indefinite (in the sense of, it is neither positive definite [making $Q>0$ always], negative definite [$Q<0$] or semi-definite [one sign + zero] but indefinite, i.e., taking both signs). For us, that means it violates CSI in some directions, and does not in others. Except if $p$ and $q$ are collinear, or if either $p$ or $q$ is zero.
The case where $\det M=0$ are interesting, and correspond to different cases, some being the local version of the ones we've been seeing already:
- $\nabla\varphi=0$, i.e., no phase structure: real fields. Then $M=gg^{\mathsf T}$ is positive semi-definite: violation in every direction except along the level curve of $\Lambda$, where $Q=0$. This is local version of the "real fields violate CSI everywhere their amplitude differs".
- $\nabla\Lambda=0$, i.e., amplitude ratio locally flat. Then $M=-pp^{\mathsf T}$ is negative semi-definite: no violation in any direction. The uniform-and-homogeneous no-go, here local but also what happens on the whole plane for LG$_{\pm\ell}$.
- $p$ and $q$ both nonzero but collinear, fully non-violating when $|\nabla\Lambda|<|\nabla\varphi|$, fully non-violating otherwise.
For the rest: no point of a rank-2 2D structured beam can be non-violating in all directions, unless the two gradients are collinear (or zero).
There is a cone of violation.
15:50 (CET).
I have given a case of everywhere-violating uniform beam, I will now give the other case of a bimodal homogeneous field violating CSI. From the previous result, this means that this field is non-uniform in space. We'll find that it has the shape of a bathtube with $G^{(1)}=2c\cosh k(x-L/2)$ over $x\in[0,L]$ (translational over $y$). The violation is everywhere on $]0,L]$ iff $k\ge2\pi/L$ (it is only somewhere otherwise).
We work on the domain $[0,L]^2$, and now consider: $$\phi_a(\mathbf{r})\equiv\sqrt c\,e^{-k(x-L/2)/2},\qquad \phi_b(\mathbf{r})\equiv\sqrt c\,e^{+k(x-L/2)/2}\,e^{2\pi ix/L}$$ also only $x$-dependent. Then, we can check that: $$\Lambda=k(x-L/2)\quad\text{and}\quad \varphi=2\pi x/L$$ and also that $$\Psi^{(2)}(\mathbf r_1,\mathbf r_2)=2c\,e^{i\pi(x_1+x_2)/L}\cosh\!\Big[(k+2\pi i/L)\tfrac{x_2-x_1}{2}\Big]$$ so that, from $|\cosh(a+ib)|^2={1\over2}\big(\cos(2b)+\cosh(2a)\big)$: $$G^{(2)}(\mathbf r_1,\mathbf r_2)=|\Psi^{(2)}(\mathbf r_1,\mathbf r_2)|^2=2c^2\big[\cosh(k\delta)+\cos(2\pi\delta/L)\big],$$ depends on $\delta\equiv x_2-x_1$ alone, and thus on $\mathbf{r}_2-\mathbf{r}_1$ alone. This is homogeneous.
Now for the violation: $G^{(2)}(0)=4c^2$ so that we get the violation $G^{(2)}(\boldsymbol\delta)>G^{(2)}(0)$ iif $\cosh(k\delta)+\cos(2\pi\delta/L)>2$. Interestingly, this looks very similar to the iff statement $\cosh\Delta\Lambda+\cos\Delta\varphi>2$. We can choose $k$ so let us define $\kappa\equiv kL/2\pi$ and also $s\equiv 2\pi\delta/L\in(0,2\pi].$ The condition for violation becomes $\cosh(\kappa s)+\cos s>2$ but since $$\cosh s+\cos s=2\sum_{n\ge0}\frac{s^{4n}}{(4n)!}=2+\frac{s^4}{12}+\frac{s^8}{10080}+\cdots$$ the we have violation of CSI for all $s$, i.e., everywhere on $]0,L]$. Reciprocally, if $\kappa<1$, there is a region of no violation, which we can see from the series expansion $\cosh(\kappa s)+\cos s-2=\frac{s^2}{2}\big(\kappa^2-1\big)+\frac{s^4}{24}\big(\kappa^4+1\big)+O(s^6)$. The leading term is strictly negative, so there exists $s_0>0$ with $\cosh(\kappa s_0)+\cos s_0<2$ on $[0,s_0]$, i.e., a whole interval around the diagonal where $G^{(2)}(\delta)<G^{(2)}(0)$, i.e., CSI are satisfied (and exhibiting bunching).
This thus provides quite a flexible homogeneous beam that can violate wholly or partly the CSI.
18:14 (CET).
Note that for homogeneous fields, antibunching, i.e., $G^{(2)}(\boldsymbol\delta)>G^{(2)}(0)$, is equivalent to CSI violation $\big[G^{(2)}(\mathbf r_1,\mathbf r_2)\big]^2>G^{(2)}(\mathbf r_1,\mathbf r_1)\,G^{(2)}(\mathbf r_2,\mathbf r_2)$, which is why they insist on homogeneity. Their criterion in this case also yields $\cosh\Delta\Lambda+\cos\Delta\varphi>2$ directly which remains true in our (more) general case.
17:19 (CET).
One cannot have CSI violation for a bi-modal field that is both uniform and homogeneous.
This is one or the other.
Homogeneity, i.e., $G^{(2)}(\mathbf r_1,\mathbf r_2)$ depending only on $\mathbf{r}_1-\mathbf{r}_2$ alone, forces $\rho_a(\mathbf{r})\rho_b(\mathbf{r})=c$ a constant where we remind $\phi_{a/b}=\rho_{a/b}(\mathbf{r})e^{i\theta_{a/b}}$. Indeed, since $\Psi^{(2)}(\mathbf r,\mathbf r)=2\phi_a(\mathbf r)\phi_b(\mathbf r)$, then $G^{(2)}(\mathbf r,\mathbf r)=|\Psi^{(2)}(\mathbf r,\mathbf r)|^2=4\rho_a^2\rho_b^2$ and since $G^{(2)}(\mathbf{r}_1,\mathbf{r}_2)=F(\mathbf{r}_1-\mathbf{r}_2)$, then $G^{(2)}(\mathbf{r},\mathbf{r})=F(\mathbf{0})$ pinning the value of $\rho_a\rho_b$ (to the square root of this over 4).
On the other hand, uniformity, pins $\rho_a^2+\rho_b^2=\text{const}$. But sum and product both constant means $\rho_a$ and $\rho_b$ are themselves constant, so uniformity+homogeneity$\implies\Delta\Lambda=0$. And this invalidates the condition for CSI violations. QED.
16:44 (CET).
So, now, Claude can help us find an explicit case:
Let us take as our two modes: $$\phi_a(x,y)=\frac{\sqrt2}{\sqrt{1+e^{2\Lambda(x)}}},\qquad \phi_b(x,y)=\phi_a(x,y)\;e^{(1+i)\Lambda(x)}$$ where $\Lambda$, we remind, is $\Lambda(\mathbf{r})\equiv\ln\big|\phi_b(\mathbf{r})/\phi_a(\mathbf{r})\big|$ so here we are making a self-referential definition of sort and need to ensure that the log of the modulus of their ratio spits the $\Lambda$ out, and you can see that it does (the complex exponential goes with the modulus and the log sees $e^\Lambda$). Note that there's only $x$ dependence in the amplitude, but we're looking for a uniform field anyway—and indeed $|\phi_a|^2+|\phi_b|^2=2$ everywhere—so invariance of $y$ shouldn't worry us. Now for $\Lambda$, Claude suggests: $$\Lambda(x)\equiv\operatorname{arcsinh}\!\big((2x-1)\sinh\pi\big)\,.$$ The arcsinh is chosen to ensure orthonormality of the modes, since, using $e^{\Lambda}/(1+e^{2\Lambda})=1/(2\cosh\Lambda)$, $$ \int_0^1\!\!\int_0^1\phi_a^*\phi_b\,dx\,dy=\int_0^1\frac{dx}{\cosh\Lambda(x)}e^{i\Lambda(x)}=\frac{1}{2\sinh\pi}\int_{-\pi}^{\pi}e^{i\Lambda}\,d\Lambda=0\,.$$
Since $\Lambda$ is real, $\theta_a=0$ and $\theta_b=\Lambda$ so that $\Delta\theta=\Lambda$ and that's the 1-Lipshitz criterion of our previous result. Therefore CSI are violated everywhere except on the set $\{\Lambda(\mathbf r_1)=\Lambda(\mathbf r_2)\}$ but this is simply the set where $x_1=x_2$, which is a cube in the hypervolume of dimension 4, and a cube has four-dimensional area zero.
Therefore, we have a two-mode uniform field which violate CSI almost everywhere (in all cases, except where the CS inequalities saturate by construction, but this is both unavoidable and this also has measure zero).
The sum of the two modes is what is uniform. Each varies in $x$ (not in $y$) and they compensate. This is also something true at the plane where the fields are thus defined. In a beam, they will distort, not being eigenmodes of the paraxial propagator.
Note that we don't even need to refer to Lipschitzity in this explicit case, we can also check directly that since both $\Delta\Lambda$ and $\Delta\varphi$ are the same $\Lambda(x_2)-\Lambda(x_1)$, the iff criterion reads
$$ \sinh^2\!\left(\frac{\Delta\Lambda}{2}\right)>\sin^2\!\left(\frac{\Delta\Lambda}{2}\right) $$
which, for $\Delta\Lambda\neq0$, is always true since $\sinh|x|>|x|\ge|\sin x|$.
22:24 (CET).
I was just discussing how to get violation everywhere in a homogeneous (constant intensity) beam, providing cases of violations everywhere but on restricted (open) domains. To have it everywhere‒everywhere... follow me:
I remind that violation requires $\sinh^2(\Delta\Lambda/2)>\sin^2(\Delta\varphi/2)$. If $\Delta\Lambda=0$, the lhs is zero and nothing can be violated. So one needs $\Lambda(\mathbf r_1)\neq\Lambda(\mathbf r_2)$ for every pair of distinct points, i.e., $\Lambda$ should be injective on the domain.
But, there is no $\Omega\subset\mathbb R^2\to\mathbb R$ continuous injection.
Therefore rank 2 (bimodal) states in 2D can never violate CSI strictly everywhere.
This sounds like a result but on the diagonal, the CSI are always saturated, so the good criterion is of course "violated everywhere it could".
And this turns out to be easy to achieve.
Since $|\Delta\varphi|<|\Delta\Lambda|$ is a sufficient condition for violation, let us assume a 1-Lipschitz function $g$ of $\Lambda$, such that $\varphi=g(\Lambda)$ with $|g'|\le1$. This ensures, by definition, that $|\Delta\varphi|\le|\Delta\Lambda|$ for all $\boldsymbol{r}$. It's not strict, otherwise we'd be set, but since $x\le y<\pi\implies\sin(x)\le\sin(y)<\sinh^2(y)$ then as long as $|\Delta\Lambda|<\pi$, we have the iff condition for CSI violation.
So now let's turn to the case $|\Delta\Lambda|\ge\pi$: this one is actually even more straightforward as we don't need Lipschitzity, since we always have $\sinh^2(\Delta\Lambda/2)\ge\sinh^2(\pi/2)\approx5.23>1\ge\sin^2(\Delta\varphi/2)$ and therefore CSI are violated again. Therefore:
For $\varphi=g(\Lambda)$ a 1-Lipschitz function, the CSI are violated everywhere except at the set $\{\Lambda(\mathbf r_1)=\Lambda(\mathbf r_2)\}$ which includes at least $\boldsymbol{r}_1=\boldsymbol{r}_2$ and possibly other points but that can be of measure zero almost everywhere.
There is a bit of work to show that the measure is zero:
Let us define $E\equiv\{(\mathbf r_1,\mathbf r_2)\in\Omega\times\Omega:\Lambda(\mathbf r_1)=\Lambda(\mathbf r_2)\}$ the set of CSI saturation. It has measure: $$|E|=\int_\Omega\Big|\{\mathbf r_2:\Lambda(\mathbf r_2)=\Lambda(\mathbf r_1)\}\Big|\,d\mathbf r_1$$ From there, measure theory and calculus establish that $$|E|=\int_\Omega\big|\Lambda^{-1}(\Lambda(\mathbf r_1))\big|\,d\mathbf r_1=\int_\Omega\nu\big(\{\Lambda(\mathbf r_1)\}\big)\,d\mathbf r_1=\int_{\mathbb R}\nu(\{c\})\,d\nu(c)=\sum_{c\ \text{atom}}\nu(\{c\})^2$$ where $\nu$ is the pushforward of area measure by $\Lambda$, i.e., for a set of values $B\subseteq\mathbb R$,
$$ \nu(B)\equiv\big|\Lambda^{-1}(B)\big|=\big|\{\mathbf r\in\Omega:\Lambda(\mathbf r)\in B\}\big| $$
i.e., how much area does a $\Lambda$-value in $B carries in space. An atom of $\nu$ is a single value $c$ with $\nu(\{c\})>0$, i.e., a value which is taken by the function on a set of nonzero measure. The integrand $\nu(\{c\})$ is nonzero only at atoms, and atoms are countable. Therefore, we have established that:
$∣E∣=0$ iff $\nu$ has no atoms, i.e., iff no value of $\Lambda$ is taken on a set of nonzero measure.
Or, equivalently:
$$\boxed{|E|=0\iff \text{no level set of }\Lambda\text{ has positive area}}\,.$$
In such a case, CSI are violated almost everywhere. Now we just need to find a case with uniform intensity.
19:15 (CET).
In my ℤ post on structured vs homogeneous (I'm not copying all of them here), __Y_Y__ I was just discussing how to get violation everywhere in a homogeneous (constant intensity) beam, providing cases of violations everywhere but on restricted (open) domains. To have it everywhere‒everywhere... follow me:
I remind that violation requires $\sinh^2(\Delta\Lambda/2)>\sin^2(\Delta\varphi/2)$. If $\Delta\Lambda=0$, the lhs is zero and nothing can be violated. So one needs $\Lambda(\mathbf r_1)\neq\Lambda(\mathbf r_2)$ for every pair of distinct points, i.e., $\Lambda$ should be injective on the domain.
But, there is no $\Omega\subset\mathbb R^2\to\mathbb R$ continuous injection.
Therefore rank 2 (bimodal) states in 2D can never violate CSI strictly everywhere.
This sounds like a result but on the diagonal, the CSI are always saturated, so the good criterion is of course "violated everywhere it could".
And this turns out to be easy to achieve.
Since $|\Delta\varphi|<|\Delta\Lambda|$ is a sufficient condition for violation, let us assume a 1-Lipschitz function $g$ of $\Lambda$, such that $\varphi=g(\Lambda)$ with $|g'|\le1$. This ensures, by definition, that $|\Delta\varphi|\le|\Delta\Lambda|$ for all $\boldsymbol{r}$. It's not strict, otherwise we'd be set, but since $x\le y<\pi\implies\sin(x)\le\sin(y)<\sinh^2(y)$ then as long as $|\Delta\Lambda|<\pi$, we have the iff condition for CSI violation.
So now let's turn to the case $|\Delta\Lambda|\ge\pi$: this one is actually even more straightforward as we don't need Lipschitzity, since we always have $\sinh^2(\Delta\Lambda/2)\ge\sinh^2(\pi/2)\approx5.23>1\ge\sin^2(\Delta\varphi/2)$ and therefore CSI are violated again. Therefore:
For $\varphi=g(\Lambda)$ a 1-Lipschitz function, the CSI are violated everywhere except at the set $\{\Lambda(\mathbf r_1)=\Lambda(\mathbf r_2)\}$ which includes at least $\boldsymbol{r}_1=\boldsymbol{r}_2$ and possibly other points but of measure almost zero.
There is a bit of work to show that the measure is zero:
Let us define $E\equiv\{(\mathbf r_1,\mathbf r_2)\in\Omega\times\Omega:\Lambda(\mathbf r_1)=\Lambda(\mathbf r_2)\}$ the set of CSI saturation. It has measure: $$|E|=\int_\Omega\Big|\{\mathbf r_2:\Lambda(\mathbf r_2)=\Lambda(\mathbf r_1)\}\Big|\,d\mathbf r_1$$ From there, measure theory and calculus establish that $$|E|=\int_\Omega\big|\Lambda^{-1}(\Lambda(\mathbf r_1))\big|\,d\mathbf r_1=\int_\Omega\nu\big(\{\Lambda(\mathbf r_1)\}\big)\,d\mathbf r_1=\int_{\mathbb R}\nu(\{c\})\,d\nu(c)=\sum_{c\ \text{atom}}\nu(\{c\})^2$$ where $\nu$ is the pushforward of area measure by $\Lambda$, i.e., for a set of values $B\subseteq\mathbb R$,
$$ \nu(B)\equiv\big|\Lambda^{-1}(B)\big|=\big|\{\mathbf r\in\Omega:\Lambda(\mathbf r)\in B\}\big| $$
i.e., how much area does a $\Lambda$-value in $B carries in space. An atom of $\nu$ is a single value $c$ with $\nu(\{c\})>0$, i.e., a value which is taken by the function on a set of nonzero measure. The integrand $\nu(\{c\})$ is nonzero only at atoms, and atoms are countable. Therefore, we have established that:
$∣E∣=0$ iff $\nu$ has no atoms, i.e., iff no value of $\Lambda$ is taken on a set of nonzero measure.
Or, equivalently:
$$\boxed{|E|=0\iff \text{no level set of }\Lambda\text{ has positive area}}\,.$$
In such a case, CSI are violated almost everywhere. Now we just need to find a case with uniform intensity.
19:12 (CET).
Now at the two-photon level—which for CSI is the one that interests us—the Monken et al.[3] transfer theorem brings the pump profile at the mean transverse coordinate: $$\Psi^{(2)}(\mathbf r_1,\mathbf r_2)\propto\mathcal W\!\left(\frac{x_1+x_2}{2},\ \frac{y_1+y_2}{2},\ Z\right)\big(\mathbf H_1\mathbf V_2+\mathbf V_1\mathbf H_2\big)\,.$$
This is also in the limit where $\Phi$ the Fourier transform of the phase-matching function is constant (thin crystals), so we can now get rid of it.
As I just said, we study the full spatial landscape while they study homogeneous beam structures. But they need to get there. Before any homogenization, the link between us and them is simply: $$G^{(2)}(\mathbf r_1,\mathbf r_2)=2\left|\mathcal W\!\left(\frac{x_1+x_2}{2},\frac{y_1+y_2}{2},Z\right)\right|^2,\qquad G^{(2)}(\mathbf r_i,\mathbf r_i)=2\left|\mathcal W(x_i,y_i,Z)\right|^2$$ And we find again the coincidence rate thus giving us the pump profile.
The two-photon wavefront is then processed to be transformed into: $$\Psi^{(2)}(\mathbf r_1,\mathbf r_2)\propto\mathcal W\!\left(\frac{x_1+x_2}{2},\ \frac{y_1-y_2}{2},\ Z\right)\big(\mathbf H_1\mathbf V_2+(-1)^\zeta\mathbf V_1\mathbf H_2\big)$$ where $\zeta=0$ for Caetano and Souto Ribeiro[4] but becomes $\zeta=1$ for Nogueira et al.[5]. In both cases, the authors have flipped the sign of one photon's $y$ coordinate (with a Dove prism and a beam-splitter, respectively). Depending on the parity of $\mathcal{W}$, this makes the function odd (Nogueira) or even (Caetano), which is compensated by the polarization, as the Bose symmetry constrains the amplitude under the combined exchange $(\mathbf r_1,\sigma_1)\leftrightarrow(\mathbf r_2,\sigma_2)$. So the antisymmetry can come from $\mathbb C^d$ and the spatial kernel is then free to be antisymmetric with the state still perfectly bosonic. Polarization can be swept out of the picture with $G^{(2)}=\sum_{\sigma_1\sigma_2}|\Psi^{(2)}_{\sigma_1\sigma_2}|^2$. The homogeneity is only along $y$, so this is only half-spatial: $$G^{(2)}(\mathbf r_1,\mathbf r_2)=2\left|\mathcal W\!\left(\frac{x_1+x_2}{2},\frac{y_1-y_2}{2},Z\right)\right|^2,\qquad G^{(2)}(\mathbf r_i,\mathbf r_i)=2\left|\mathcal W(x_i,0,Z)\right|^2$$
Both now have that at the same (transverse) positions $y_1=y_2$, their functions are identically zero. That's where their CSI violations come from, and how it gets enforced everywhere, with no spatial structure, since cross correlations will not be zero. Nogueira et al.[5] get this from their odd spatial pump profile, obtained by their $\pi$ shift of half of a Gaussian, turning it into basically a fermionic HG mode. Caetano and Souto Ribeiro[4] achieve it by blocking the pump with a wire. The two approaches are, modulo this fairly anecdotal implementation detail—a symmetry-protected zero versus a force-engineered one—fairly similar. This is homogeneous in the sense that they can displace both detectors together, and the result remains the same. In our case, the positions of the detectors determine everything. We could still, get, though, violation of CSI for homogeneous cases too in a restricted sense: violation with a uniform one-photon intensity. The CSI could be violated everywhere on a beam of constant intensity. But our degree of violation has a structure, while theirs remains constant too (in principle, infinite violation).
Concretely, in our case, $\ket{1_{+\ell}1_{-\ell}}$ and $\ket{1_c1_s}$ have the same $G^{(1)}=2R(r)^2$—the same donut structure, indistinguishable to any intensity measurement (that's my "it's the same picture" joke)—and the first never violates anywhere while the second violates over open regions with $\ell$-fold structure. That kills the objection that our violations would just be beam structure.
15:59 (CET).
The most important conceptual difference between "us" and "them" (literature of spatial SPDC correlations) is, I now believe, the Schmidt rank gap: we're in the low Schmidt rank case of studying correlations between a few (in fact, 2) modes, while SPDC people not only usually deal with Schmidt number $K\sim\left(\frac{\text{transverse extent of }\mathcal W}{\text{width of }\Phi}\right)^{2}$, typically of order hundreds or thousands, but in the homogeneous case they strive for—which they indeed put at the center of their approach—they are in the $K\to\infty$ limit. They don't have perfectly homogeneous system so if we'd try to recover them from our approach, we'd need the full multimode theory, and we're not there yet. But, the difference is now clear: we have real, spatially structured CSI violations, while they violate it everywhere... and you know what they say: if it's for everybody, then it's for nobody. So I think we can claim we are the real spatial-CSI violating people.
(Here careful about Schmidt rank [number of modes] and Schmidt number [how many modes carry the weight]; Caetano and Souto Ribeiro[4] have infinite Schmidt rank but actually fairly small Schmidts numbers)
For SPDC, the two-photon amplitude factorizes into a pump term at the mean coordinate and a phase-matching term at the relative one:
$$ \Psi^{(2)}(\mathbf r_1,\mathbf r_2)\propto\mathcal W\!\left(\frac{x_1+x_2}{2},\ \frac{y_1+y_2}{2},\ Z\right)\Phi(x_1-x_2,\ y_1-y_2;Z) $$
$\mathcal W$ is the pump field propagated to the detection plane $Z$ (although it got transformed into the two-photon beam in the process). A UV filter blocks the original pump beam after the crystal.
Diagonal and trace give different things. On the diagonal, the pump image survives intact:
$$ \Psi^{(2)}(\mathbf r,\mathbf r)\propto\mathcal W(x,y,Z)\,\Phi(0,0;Z) $$
whereas the trace is a convolution:
$$ G^{(1)}(\mathbf r)\propto\int d^2\mathbf v\;\big|\mathcal W(\mathbf r-\mathbf v,Z)\big|^2\,\big|\Phi(2\mathbf v;Z)\big|^2 $$
Let us define $a$ the transverse scale of the structure imprinted on the pump—for Caetano and Souto Ribeiro[4]'s imaged 250 µm wire, $a\approx0.5$ mm—and $\sigma$ the width of the pair-correlation factor $|\Phi|^2$ at the detection plane, i.e., how far apart the twins land. From the phase-matching angular spread, $\sigma\sim Z\sqrt{\lambda/L}$, which with $Z=0.75$ m, $\lambda=884$ nm and $L=1$ cm gives $\sigma\approx7$ mm. The convolution then leaves a residual contrast for the wire of order $a/\sigma\approx7\%$, smeared over 7 mm instead of $2a=1$mm. This is the idea of quantum imaging: the object is not visible in one-photon observables but is revealed in two-photon ones.
rank-2 beam (ours) SPDC beam (theirs) pair diagonal $\Psi^{(2)}(\mathbf r,\mathbf r)$ $2\phi_a\phi_b$, dark on a cross of nodes $\mathcal W(x,y,Z)$, the pump image, sharp intensity $G^{(1)}(\mathbf r)$ $\lvert\phi_a\rvert^2+\lvert\phi_b\rvert^2$, a doughnut $\lvert\mathcal W\rvert^2$ convolved with $\lvert\Phi\rvert^2$: flat, cm-wide link between them same two functions, exact convolution over $\sim10^2$ modes 15:08 (CET).
I was wondering about bosonization as understood by bosonizers and by exciton people. I asked Nina and Mikhail Glazov about their opinion. Misha replied on the spot, that he was in Karelia and could not attend to the matter in details right away but would try on Thursday or Friday (this was on a Tuesday). And sure enough, first thing in the morning of Thursday came a long and detailed email.
There he explained that «in 1D the classical Fermi-liquid picture for interacting fermions (with repulsive interactions: typically, one band is considered like in metal or doped semiconductor) does not work» and that «it is because fermionic quasiparticles are ill defined», but that «It turned out that the problem of interacting particles can be elegantly solved in 1D by introducing novel bosonic modes (charge density/plasma wave and spin density/spin wave).»
All this is already very relevant, but even more the following:
Note that this bosonization approach is different from “pairing” of electrons and holes to form excitons: In the excitonic case we have (exact) solution of a two-body problem, namely, the “exciton”, and we approximate the manybody state as appropriately (anti)symmetrized products of single “exciton” states. In the case of the Luttinger liquid (bosonized 1D fermions) we consider collective modes right from the start.That makes the 1D case very specific. He then observes:
Note that in 2D/3D cases the spectrum of Fermi liquid has both single-particle branch (quasielectrons/quasiholes) and manybody excitations (sound/plasmons and spin waves). In 1D case there are no single-particle excitations, at least at small wavevectors.And back to 1D:
In 1D Luttinger liquid the key problem is somewhat different although it is also related to the commutation relations. Technically, they do the following: The total density of electrons is split into the densities of left- and right-movers, and commutation relations of (the Fourier components of) these densities are shown to have specific approximate form (that can be used to express everything in terms of bosonic operators). The accuracy of fulfilment of these relations is the better, the smaller are the wavevectors [in "Bosonization for beginners — refermionization for experts" (https://arxiv.org/pdf/cond-mat/9805275) it is briefly mentioned in Sec. 5 below Eq. (34) then they introduce the scale a=1/k_F]. The applicability of this procedure is, however, unrelated to the number of bosons, etc. as long as the interactions in the initial fermionic Hamiltonian are of “density-density” form.Then he also comments on suitability of bosonization for phonons and magnons, which however isn't my point as I worry about the structural impossibility of the concept due to the compositeness of excitons, not on deviations from nonlinearities (which are dealt with correctly through perturbation, in my understanding).
He also makes more—and interesting—comments on regimes where excitonic breakdown would be relevant (small QDs—we actually have a paper on this[6]—Rydberg excitons and polaritons, etc.) and concludes with:
As a historic anecdote, I also remember Robert Suris telling me about his argument with Leonid Keldysh back in 1970s: as you remember, Keldysh took great care of keeping fermionic corrections to the exciton commutation relations when studying exciton condensation [1] while Suris argued that it is enough to consider the exciton-exciton scattering amplitude [2]. In any case, their results agreed both qualitatively and quantitatively… Incidentally, Combescot approach showed that the correct calculation of the two-exciton scattering amplitude (in her technique) yields the same result as other people obtained using exciton language.So no need of excitons in the case Haldane tackled exactly... but is it because the concept can be bypassed altogether, or because it is ill-defined? If it is bypassed, is it because they do something similar to Combescot, and tackle the full many-body, collective aspect from the start? If it is ill-defined, does it point out in this "simple" case where things can be solved exactly, that excitons are indeed problematic and cannot fit the picture? Or is it because the problems have not been connected and both communities ignore each other? I don't know about Combescot referring to Haldane (though I should check more carefully), but I know that the bosonization experts who visited us at ICMM several times 1/2/3 didn't know about her and her problems with excitons. In this case, what if one takes this exact case, as done by, say, Lee and Eric Yang[1], write an interacting bosonic Hamiltonian for this and compare the two theories. One remains exact, doesn't it? The other is the standard practice for most of the excitonic/polaritonic community. How do these relate?
11:07 (CET).