/home/andrew$

Some thoughts on physics, statistics, computing & technology

Double reading, writing and descent

September 17, 2026 — Andrew Fowlie

There are several doubles in my title, for I've been reading two books, writing two sets of teaching notes, and thinking about double descent, the phenomenon in deep learning.

On the reading, as planned I've started a more substantial novel, Hard Times, considered by some critics to be Dickens' most complete or perfect work, since it is compact and without the sprawling plots and character ensemble of Dickens' other celebrated works. I was worried it would lack warmth and affection, but so far I am thoroughly enjoying it and its connection to educational theory.

I am also reading the wonderful Mrs Frisby and the Rats of NIMH with my daughter. This is everything a children's book should be, and I will write more when we have finished it.

As for writing, I am continuing my In Your Face For Anybody series, and am partway through an introduction to chatbots, starting with ELIZA in the 1960s and ending with contemporary large language models. It is a work in progress [pdf], but I'm enjoying creating it. The series began with my notes on quantum mechanics and Bell's theorem [pdf].

As it's the beginning of term, I also dug up my data analysis booklet for our lab students [pdf]. The booklet was inspired by one given to me in my first year as an undergraduate, though it's not as detailed or in depth, or perhaps as insightful. The ones we were given in 2005 were referred to as our red books and they were a superb practical guide to data analysis in the laboratory and some frequentist statistical concepts. I think they were written by Prof. Tom Shanks, more affectionately known as Dr. Tom at that time, though it may have been a different Tom.

Last but not least: Double Descent! This is a remarkable phenomenon in data science, that at first seems rather mysterious. Consider linear regression on \(n\) data points using \(p\) basis functions with \(p\) coefficients, $$ y = \sum_{i=1}^p \theta_i f_i(x) $$ We train our model on the \(n\) data points; select a point estimate for the parameters; and evaluate performance on a hold-out data set of another, say, \(m\) data points.

What happens when we increase \(p\) starting at \(p \ll n\)? At first, the model is under-parameterised. It lacks flexibility to fit the data points; it is biased, but not wild, that is, it does not suffer from high variance. As we increase \(p\), the performance improves as the model is sufficiently flexible to fit the data. This is the first descent.

As we increase \(p\) yet further, performance deteriorates. Our model becomes too wiggly; we are no longer primarily bias-dominated, but we suffer from variance. At \(p \approx n\), we perfectly fit the training data, but performance on the test data is poor, as we have overfit. This is the interpolation limit.

In classical statistics, that would usually be the end of the story. We find the optimal balance between bias and variance, and that's that. Here comes the twist though. Keep on increasing \(p\) beyond \(n\). Something strange can happen. The system is overparameterized and underdetermined, and there are many ways of achieving a perfect fit to the training data. The performance on the hold-out data improves; this is the second descent.

As we increase the number of basis functions, we are able to create combinations of basis functions that are smooth but perfect fits through the training data. Nothing in our training, however, favoured smoothness: we fitted to the training data without any regularization term. Why, then, are these smooth solutions selected? They are selected by the implicit bias of the gradient descent optimizer in this context. Whilst there are many solutions that perfectly fit the training data, the gradient descent algorithm is biased towards the solution that minimizes the \(L_2\)-norm of the parameters. That is a smooth solution.

The long and short of it is that by increasing \(p\) we create more possibilities for smooth functions that fit the training data perfectly and generalise well. Since we are using point estimates, there is no explicit Occam-type penalty for increasing \(p\). These smooth interpolating functions, that weren't selected or available near the interpolation threshold, are selected by the implicit bias of the optimization algorithm, and being smooth functions that fit the training data, they typically generalise well.

Tags: reading, writing, double-descent

Baryogenesis, finally

September 09, 2026 — Andrew Fowlie

Sakharov's third condition for baryogenesis [A.D. Sakharov, Zh. Eksp. Teor. Fiz. Pis'ma 5, 32 (1967); JETP Lett. 91B, 24 (1967)] is usually stated as: departure from thermal equilibrium or interactions out of thermal equilibrium or some such. I've always had some trouble with the textbook arguments for it, but I think I've now had a personal epiphany.

The textbook argument, e.g., Trodden's notes, considers the expected baryon number in equilibrium: $$ \langle B \rangle = \text{Tr}\left(e^{-\beta H} B\right) $$ Inserting an identity, \( (\text{CPT}) (\text{CPT})^{-1} \), using the cyclicity of the trace and the invariance of the Hamiltonian under CPT: $$ \begin{align} \langle B \rangle =& \text{Tr}\left((\text{CPT}) (\text{CPT})^{-1} e^{-\beta H} B\right)\\ =& \text{Tr}\left(e^{-\beta H} (\text{CPT})^{-1} B (\text{CPT}) \right)\\ =& - \text{Tr}\left(e^{-\beta H} B\right)\\ =& -\langle B \rangle \end{align} $$ and thus \( \langle B \rangle = 0 \) in equilibrium.

The thing that confused me was that it seemed that we almost begged the question. We started by assuming that \(H\) was the only relevant macroscopic parameter. We didn't include any chemical potential for baryon number in our ensemble. Suppose we took: $$ \langle B \rangle = \text{Tr}\left(e^{-\beta H + \mu_B B} B\right) $$ The partition function now no longer commutes with CPT and we cannot conclude that \(\langle B \rangle = 0\). Try it. The above argument won't go through. Can we do that? I don't think so. Sakharov's first condition for baryogenesis was that \(B\) was not conserved. In which case, there is no chemical potential in the equilibrium ensemble.

But what if we can give a chemical potential to some other conserved quantity? What? Well, \(B - L\) is a natural candidate, as unlike individual baryon and lepton number, it is conserved. In that case, we could write the expectation of baryon number as, $$ \langle B \rangle = \text{Tr}\left(e^{-\beta H + \mu_{B - L} (B - L)} B\right) $$ Because of the chemical potential, we cannot conclude that \(\langle B \rangle = 0\). Is this a loophole to Sakharov's conditions? No, we've in fact rediscovered leptogenesis! We give a chemical potential for \(B - L\) and equilibration redistributes the charges such that \( \langle B \rangle \neq 0\). We convert a \(B- L\) asymmetry to a baryon asymmetry.

Did this require a departure from thermal equilibrium? Well, under standard cosmology, there are no chemical potentials after reheating, as the inflaton decays into a hot thermal bath with no chemical potentials. Creating a chemical potential for \(B-L\) from scratch requires an out-of-equilibrium process. Thus Sakharov's third condition holds.

What if we don't assume that there are no chemical potentials after inflation? Maybe inflation can produce a chemical potential for \(B-L\)? Is that a loophole? No, we've in fact discovered inflationary baryogenesis and changed the era in which baryogenesis occurs.

What if we don't assume inflation at all? In that case, indeed, we can input what we like as an initial condition, e.g., an equilibrium state with non-zero baryon number.

Tags: baryogenesis, statistical-mechanics, physics

AI & Misogyny

September 08, 2026 — Andrew Fowlie

Read The New Age of Sexism: How the AI Revolution is Reinventing Misogyny by Laura Bates. This was an urgent and alarming book, that occasionally trod close to but avoided alarmism, about the interplay and impacts of AI and technology on sexism and discrimination.

The book at times read like a thriller: it was emotive, shocking and fast-paced, but wasn't particularly theoretical, either on technology and AI; on the causes and origins of sexism; or on our economic system and the rise of big tech and monopoly capitalism. The book works, however, as an eye-opener about where we are already headed and as a passionate call to arms to think and do something about it.

Instead of sacrificing anything and everything to the whims of men like Mark Zuckerberg, Elon Musk, Phillipp-Fussenegger, Cameron-James Wilson and a thousand others like them, instead of breathlessly enabling their ruthless pursuit of financial profit in the name of progress, we need to start from a place of sustainability, of fairness, of real, social progress, not technological development for the sake of it ...

... a different baseline of what progress looks like, what its goals are, who it serves and how it happens.

Tags: reading, ai, society

17 papers, 2.6 sigmas, and 1 event

September 04, 2026 — Andrew Fowlie

The LZ experiment recently announced (1 September) a new result through a talk and preprint paper.

The paper documents something strange: a single, unexplained event that could be consistent with an interaction with a Dark Matter particle. The data were analysed using frequentist, and specifically Fisherian, statistical methods described here. In particular, evidence was quantified using a \(p\)-value, that is, the probability of obtaining data that were at least as extreme as that observed given the background-only (i.e., no Dark Matter) hypothesis. The \(p\)-value was about 0.5%. For historical reasons, this is communicated as \(2.6\sigma\), as 0.5% corresponds to the probability mass contained in the tail of a Normal distribution at \(Z = 2.6\).

Hey ho, you might think. That ain't much evidence for an extraordinary claim, especially since it is contingent on their background and detector modeling to the 1 rogue-event level, and given all we know about \(p\)-values, the replication crisis, and the strength of evidence. That said, it might just about be below a threshold proposed by Benjamin et al. (2017) in social and medical sciences, thought not close to the physics typical requirement of \(5\sigma\) for a discovery.

In the two days that followed the paper, however, there were 17 theory papers that presented particular models of Dark Matter that could explain that 1 event. Hmm. I haven't read them well enough to comment on how well-motivated these explanations are, though the most popular appears to be a Higgsino explanation (see e.g., Fan & Reece), where a supersymmetric partner of the Higgs plays the role of Dark Matter. However, a paper today suggests Higgsinos may be in tension with so-called sideband results from LZ. That means, they would predict Dark Matter signals in other searches performed by LZ, but none were seen.

There was a particular statement in Fan & Reece that caught my eye:

The look-elsewhere correction depends on the total number of models that were fit to the data, and as such is arguably too conservative when, as we will argue, there is one particularly compelling model to focus on
Hmm. One of the infamous things about \(p\)-values is that they depend on the experimentalists' analysis plan. That is, they depend on what the experimentalists would have done, were the data different. Part of computing them correctly thus involves correcting for the look-elsewhere effect (LEE). This means, correcting for all the things you, the experimenter, would have looked at had the data been different. There is a famous reductio ad absurdum about a statistician and electrician, by Edwards and Pratt, described here that shows how strange this can be.

The long-and-short of this is that the \(p\)-value depends on the experimentalists' analysis plan, which should ideally be declared before data collection. I don't think it's reasonable to claim after the fact that
Ah! But I was only interested in this subset of tests anyway (which happened to include tests that produce the smallest \(p\)-values). Let's just use those ones. Thus the \(p\)-value is smaller!
This would have been perfectly fine pre-data. I would have no objection to pre-data, or pre-unblinding of data, someone saying, let's use an analysis plan that only performs these tests to maximise the statistical power for these particular Dark Matter models that we find plausible. Post-data, however, it's too late, at least within a frequentist paradigm.

On the other hand, perhaps Fan & Reece are subconsciously thinking as Bayesians. In Bayesian frameworks, we don't need to consider the sampling plan or what we'd have done were the data different: our conclusions depend only on the data we did observe. This is connected to the likelihood principle. In that case, perhaps a more charitable interpretation is that:

Ah! I don't care about your error rates and what you'd have done were data different. I just care about what I should believe given the data at hand. The Dark Matter models that explain the event were a priori plausible to me. Given that, I am more inclined to believe that they explain the data.
I think that's a defensible point of view.

Tags: dark-matter, statistics, physics

Tall Tales and Wee Stories

September 01, 2026 — Andrew Fowlie

Read Tall Tales and Wee Stories, a collection of Billy Connolly's stand-up material. I chose it to distract myself from other matters going on, and it mostly did the trick. I didn't know many of Connolly's routines, though remember his voice and manner from TV and from some VHS tapes my parents had in the 90s.

The routines are often funny, though occasionally misogynistic to a modern reader. Obviously something is lost on the page, especially the chaos, spontaneity, and physicality of his live performance, such as the one about getting drunk from the feet upwards. I was surprised by the lack of reflection, nostalgia, politics or social commentary in his material: I presumed he'd draw on his years as a welder or his working-class upbringing, but the routines are somewhat apolitical.

I'm part way through another AI/tech/society book, and then I plan to tackle a classic: Dickens or Hardy, something like that. Connolly's book was amusing, but didn't feel altogether rewarding.

Tags: reading, comedy

Dreaming is for free

August 26, 2026 — Andrew Fowlie

Read a heartbreaking story about a landslide at landfill site that killed 30 people in Guinea. Was appalled to realize that these occurrences are so common in the global south that they have their own Wikipedia entry. What an awful metaphor for the world we've created. Don't let anyone tell you a better world isn't possible.

Tags: world, tragedy

WALL-E in rep

August 26, 2026 — Andrew Fowlie

Watched WALL-E at a cinema in Qingdao, which is the center of the Chinese film industry, similar to Hollywood in the US. Chose it as it was in English and family friendly. What with the spate of remakes, I wasn't sure if it was a new version of an old film. However, as they were also showing Shawshank Redemption, it didn't come as a surprise that it was in rep.

I was pleasantly surprised by WALL-E and found it alarmingly prescient (2008). A monopolistic company, 'Buy N Large', de facto controls the world and trashes it. Robots are left behind to clean up the mess while humans escape in a spaceship. The humans become infantilised and alienated from each other, as their lives are organized by a system of robots that control the ship. The humans degenerate both physically and mentally.

In the end, humanity is saved by a robot, WALL-E, who finds beauty in a single weed growing amongst trash on Earth, and love in a fellow robot. I enjoyed the themes of endstage capitalism, environmental destruction, technology leading to alienation and decay, and the unusual resolution that it was a robot that rekindles what it means to be human. I didn't understand the lack of cultural and ethnic diversity amongst the humans on the spaceship: were only English-speaking white people saved?

Tags: sci-fi, cinema

The Machine Stops

August 21, 2026 — Andrew Fowlie

Read The Machine Stops by E. M. Forster. I hadn't previously associated him with science fiction, so it was a surprise that he wrote this book. However, the book focusses on human connections and alienation and thus perhaps isn't dissimilar thematically from his other works. The book imagines a future where we live our lives through machines, including communication and experiences. I found the Machine's preference for $n$-th hand information intriguing: first-hand and second-hand sources were considered partisan and unreliable. The society thus prefers depersonalised, smooth, generic lectures and historical accounts. Any of this sound familiar?

Tags: reading, ai, sci-fi