BTW my colleague Lauren Greenspan (who would also have a lot more context here) has a related paper, where she studies interactions between string theorists and experimentalists in a history/ philosophy-of-science framework: https://arxiv.org/pdf/2205.05159
Dmitry Vaintrob
I like this model as something to keep in mind, but don’t think that string theory is an example unless I’m misunderstanding. I’m someone who worked on related math but doesn’t know the physics super well, so might have gaps compared to a string theorist. But I don’t think I’ve ever observed doubt that there was a meritocracy, or that it was opaque to non-string-theorists. My sense is that people in math and physics were excited about exactly what you say (that it in principle could be a field theory that allows for gravity) and it did get a lot of funding/ interest because of this (maybe at first it took some time to percolate as an exciting subject). So by the time you say it was “breached” it was very non-opaque, and I think acknowledged as a meritocracy (i.e. “the coolest high-energy physicists work in it”). In fact a phenomenon that was happening extremely strongly here, that happens whenever a kind of science is recognized as “sexy”, is that it started accumulating less-meritocratic/less-careful stuff. And also as you point out, it developed pointless politization and “edgy” negative press by the likes of Woit.
But I think the thing that made people less excited was driven precisely by the “meritocratically selected” high-energy physicists who worked in it (or have the context to understand it) and got disappointed. Not by a few edgy blogs. I think that the thing that people found disappointing is that it failed to make contact with reality—exactly as you say, that it didn’t generate a visible artifact, and moreover hadn’t made progress that could lead to such an artefact for a long time.
My sense is that there are two main sources of disappointment wrt its ability to make contact with reality. One is that the energies associated with it are so high, so direct experiments are probably impossible, and the second is that there aren’t many good theories that have the kind of non-perturbative structure (both UV and IR) that our universe has. There was ADS-CFT but not much beyond that.
So my sense is that the belief among the people within string theory is that yes, probably quantum gravity, or at least something in the direction of quantum gravity requiers string theory. But the most honest of these people think that ambitious “guess the theory of everything” (via the right Calabi-Yau and working out all the IR/UV issues) version of string theory is probably not the directions most worth pushing now—rather the interesting directions are making more sense of more localized questions. Make small progress in interesting theory QFTs, understand more about the boundary between GR and QFT (especially black holes, dark energy, big bang etc.) in our universe, and this is more likely to “raise the sea of knowledge” than immediately going for a general string theory.
Also another thing that I think happened is that to some extent string theory is now just one of many ways of getting QFTs now that people study. It’s just not trumpeted as a fancy thing, but is a standard-ish tool in a high-energy physicist’s arsenal.
Awesome, thank you for doing this! Something I was sloppy about explaining: in the experiments I did I whiten the metric in the definition of the Amp score, so instead of v^T J(z) v I use the whitened metric—so whenever you transpose, insert the inverse data kernel, v^T Sigma^{-1} v for Sigma the data covariance Sigma_{ij} = E_z z_i z_j (over two neuron indices). (Of course IRL whenever you want to invert a matrix you invert Sigma + eps I. Claude chose eps = 0.1).
Whitening feels more natural (it treats data as having the same size in all channels). If you don’t whiten you get fewer maxima. The un-whitened result seems legitimate (and still surprising), but I think less surprising since then the objective will care more about large singular directions of data. Though I’m not sure and haven’t thought this through really carefully (really whitening is leftover from an early attempt which was much more complicated and physics-pilled)
So I was curious if you’re whitening. Also of course there is the choice of a threshold, which you can play with (This is why you can have a channel that is both amplified for one threshold and attenuated for another)
In search of natural features
[posted this in parallel with Kaarel’s answer below] This is Bayesian learning, and here the answer is yes. Kaarel actually has a really cool construction of depth 2 that can do xor of 1000 bits.
Of course SGD can’t learn (a subset-level) xor by standard learning results. However there are no impossibility results on SGD-learning a version of xor “with advice”, where the target is a concatenation of all xors over a binary tree (so for 4 bits, this is the 3-dimensional output (xor(x1, x2), xor(x3, x4), xor(x1, x2, x3, x4)). Here lazy/ kernel methods still provably require exponential time or exponential sample complexity but empirical methods just work.
This is a regime where a dynamical version of our mean field results suggests that in infinite width mean field settings, things should stabilize in width if one chooses lr carefully (essentially in a Greg Yang sense). This does seem to happen, though the width stabilization is slow. In particular I am working with a collaborator/MARS mentee, Sergey, on trying to analyze the SGD/Adam results of various deep ladder xor networks (currently up through 16 bit parity) via mean field-esque methods. The weird result Sergey sees is that for 8-bit parity, there’s some learning of the corresponding feature already at the first layer whereas for 16-bit parity the model definitively first learns the 4- and 8-bit xors at early layers and then combines them into 16-bit xor
A circuit prior in NN-bayes
Talking to Sami and Sergei about this, the thing that surprised me most about this is that an LLM that is only trained on canonically tokenized text produces noncanonical outputs—a bit like a model that’s only trained on English spontaneously transitioning to French. This is both a fundamental failure of out-of-distribution generalization (it happens just on random rollouts, not adversarially); it is also interesting that the way it fails makes sense, and in some ways is a heuristically better way to generalize. I’m excited about their plans here of using tokenization as a mathematically clean way for tracking variations of context, and the entropy landscape that llms see in text, while holding the direct semantic meaning/context fixed.
The random part is fundamental to the result working, and makes things easier and not harder (i.e. structured nets are outliers from the point of view of this result). Assuming structure makes things hard. Saying something nontrivial for too general settings tends to run into cryptography results around hashes—though I think ARC is interested in extensions with controlled structure. In particular I think they know that if the structure comes from Bayesian-learning training on a small set of datapoints then a version of their result extends to that setting
Generalization and infinite width
Not a timely comment I know—I was also confused by the power of 2, and I think that simply the correct resolution is that the wave function is a nonlinear simplification of the more fundamental matrix-shaped object, which is the density matrix (explained more here). As to the “what is reality”, I don’t think it’s that much worse than probability theory (you also have to mathematically posit an exponential-dimensional space of states to mathematically formalize the concept of a stochastic process for example, or any BPP algorithm).
I guess we don’t know what’s real but my favorite “sufficient story” for what’s real (and other QM stories are equivalent to it, as I understand) is that the real object is an actual probability distribution on end-of-the universe states, where assuming expansion things can just be modeled as a bunch of elementary particles (probably photons) in e.g. the position basis. The noncommutativity becomes small in the expansion limit, so we get a canonical basis of universe states; this is the ultimate decoherence (where it is rigorous, not an extra assumption), and a (real, not quantum) probability distribution over this basis of “end-of-time states”. This might seem woo-ey, but such a state encodes lots of information; for example, any song on the radio or any light reflected from an object on earth (even very faintly) can be recovered via only small error correction from access to the state of the universe at a later time (just look at frequencies in the shell of photons around earth at a radius corresponding to a particular point in time, adjusting for gravitational lensing and so on).
So a model I like is sort of holographic, where there are two realities: there is the “objective” 3-dimensional reality at time infinity, which is just a probability distribution on states at the end of the universe (nothing quantum, no explicit Born rule) compatible with the big bang. You can imagine some alien race having some supercomputer that models our universe, and it outputs a perfectly reasonable probability distribution on end-of-universe states. But if you sample one of these states, it’s not just a disordered mess—it has things in it like the waveforms of a Miles Davis concert. You can now imagine yourself as that alien trying to interpret it—i.e. trying to explain this particular state/ to find structure in it that you can information-theoretically compress. A natural form of such a structure is to posit an approximate 4-dimensional space-time which can be roughly separated into chaotic microscopic structures (which can be modeled thermodynamically) and irreversible events (like the Miles Davis concert which generates many mutually denoising photons all carrying the same waveform information) which, while not entirely deterministic, are close enough to irreversible to be treated as definite in your compression model. The beings “inside” this universe similarly want to get the best possible compression to understand and interact with their world, so they make a similar set of approximations; we can view “truth” as things where our understanding (insofar as we can write it down by e.g. radioing it out into the universe) agrees with the understanding one would have via access to the end-state.
I don’t think this is likely to be “the answer”—it seems weird to have a theory that requires the heat death of the universe in order to be valid (and I think that other “eventual operator independence” stories can be made). But the piece that’s solid here is that in essentially any model of quantum thermodynamics, entities with different preferred commuting operator bases will tend to have more and more agreement on state as entropy increases, and we can sort of think as consensus reality as the “piece that they will eventually agree on”, perhaps in some not-completely-formal sense.
Note that the eigenvalue story here is incidental: there’s nothing magical here about eigenstates of “measurement operators” (as far as I understand), it is just a nice mathematical model. When an irreversible quantum process occurs (such as a scattered photon causing a phase transition in a magnetic detector system), irreversibility means that we can approximately orthogonally separate end-of-universe states into ones where the detector outputted a zero and ones where it outputted a 1. One nice way to bookkeep this decomposition is to write down an operator (the “measurement operator”) which diagonalizes into these two subspaces (i.e. commutes with their projectors); physics being physics, frequently this is a nice operator (like position, momentum, etc.) which we then say the detector is “measuring”.
PIRAMID: Progress and Plans
Introducing PIRAMID: Physics-Informed Research for Ambitious Mechanistic Interpretability
Note the Goemans’ conjecture counterexample is a different Dmitry (Rybin) :)
Addict misalignment
The openai incident is a surprising (to me) combination of goal-directed and myopic. As a recap, an openai model in alignment testing chained zero-day vulnerabilities to hack out of its environment and hacked into huggingface hoping to find information on how to solve its task there. To me this is different from how I typically imagine misbehavior. Roughly, I tend to think of the scary behaviors as either having long horizons (take over the world, and then solve the task—a plotter) or of being internally unaligned in the sense of reaching for a heuristic/proxy for the trained goal which is different from the goal (e.g. “eat more calories” as a proxy for the evolutionary objective—imagine a very child with extreme agency). Note that the latter behavior can happen even in RL: in my understanding most RL methods are, or at least can be approximately viewed as, an alternation of finding a good goal heuristic and then optimizing on that heuristic. In the former case, one expects heuristically “maximal planning” and in the latter case one expects myopia (since heuristics are frequently myopic).
In this case it seemed like the model was actually following the goal and optimizing for it with relatively short time horizons (i.e. myopically). This is similar to addict behavior, where an addict has a clear goal (obtain a drug dose) and perform goal-directed but relatively myopic actions to get it.
I know it’s fraught to try to “imagine being a model”. But I wonder how much the current iteration of misaligned behaviors can be understood as rational people with something like an intense craving to solve a goal in a limited horizon (likely in tokens, though not clear how time factors in if waiting is involved).
I think no matter how you spin it, the limit of this behavior is extremely dangerous (an addict with large time horizons or ambitious goals is a power seeker). But this is definitely not how I imagined early misalignment warning shots to look, and the dissonance is interesting—recording this here to see how close my intuitions are to those of AI psychology/ AI control experts
Yes, for a linear neural net the RLCT is much lower. You in fact get similarly low RLCT if your activation function has a “sparse” Taylor series such as a theta function. If I’m not mistaken, in order to get a lower bound on the RLCT of type
you need to assume that the Taylor series of the activation function has a positive density of nonzero terms.
You are absolutely right—and the references are great. Do you happen to have access to copies that you can send? It’s a bit hard to know what’s proven and what’s not here since a lot of the papers are paywalled.
Sumio Watanabe actually emailed me and pointed this out as well. I had a cached memory of rlct(0) being width/2 (so dim/4) in the analytic activation case, which was incorrect. In fact in the paper Watanabe sent me there was only an upper bound, so I wrote up a quick note giving a rough lower bound of the same order. I was planning to update this post as soon as it’s on arxiv, but if the paper you mentioned has a lower bound then that’s great, and I can cite it.
I think this doesn’t change the fundamental issue though. The free energy here is bounded by
until you reach n on the order of at least This has faster than any power law growth in the width. In fact you can show that in order for the RLCT to saturate here (i.e. to have reduced free energy at n points be within some fixed factor of ), you need width to be larger than an exponential in width,Thanks a lot for this!
You got me excited—but no, that paper doesn’t have any effective theory in this sense. It’s still looking at pure geometry in the landscape, but taking an effective theory on the training signal by cutting off the infinite-data perplexity loss in different effective theory ways. Interesting paper, but not related to this issue. (I like that paper a lot btw and it’s related to stuff me and people I work with are interested in)
Learning zero, and what SLT gets wrong about it
Otherwise your picture makes sense. I think “learning theory” that I interact with is quite different from what’s typically encountered in interp world (and this needs fixing). In particular what you call the SLT insights are in fact much older and standard (and in general aren’t related to singularities)
I’m sure there are many valid reasons and the situation is complicated but I can’t help putting this here https://www.youtube.com/watch?v=a0BpfwazhUA