Personal scale good and evil is the wrong lens for a company. Brands are not your enemies or your friends, and I neither love nor hate the red blood cells of the world. The question is whether the persistence of these systems of raising money and allocating compute and human talent produce good or bad structural forces re: the ai race. In probabilistic terms, what does P(destructive ai race | lab X exists) look like compared to P(destructive ai race | lab X has shut down) look like?
testingthewaters
if you stop building towards superintelligence the chance for extinction goes up.
I think this is a very uncertain thing: https://www.lesswrong.com/posts/CEwmQWjBwaF5tai2T/even-if-others-are-less-responsible-you-can-still-make
Even if others are less responsible you can still make things worse
A few ways staying in a technology development race still makes things worse, even when there are less responsible actors around:
Creating a sense of artificial urgency, which causes people who are competitive and results driven to work harder and cut more corners
Providing a guiding star on which technical directions are promising (e.g. reasoning models, CoT, RLHF)
Legitimising irresponsible behaviour as valid response to strong competition (“it’s a competitive market, this was bound to happen”)
Entangling the motivations and incentives of the “more responsible” parties with the success of the technology in question, making any warnings seem weak and hypocritical (“if it’s so bad, how come you’re still doing it”?)
Providing an adversary or target to aim for/surpass (similar to artificial urgency)
Diffusing media/governance/public attention away from the “more irresponsible” parties, since it’s now a multi horse race
Driving hype, attention, and funding to the field via your credibility and achievements
Becoming trapped in a “we must win” mentality that causes you yourself to cut corners
Thanks for writing down and making concrete a lot of intuitions I had about the linear representation and circuit hypotheses being “slightly off” in some way. This article may be interesting, or at least the linked paper: https://picower.mit.edu/news/cognition-and-consciousness-arise-analog-computations-says-new-theory
Yeah this is more like a parody or extreme rendition, but a lot of these trends i do see when I try to get models to do something and don’t have perfect vision of what the thing will be. And of course critical situations/important alignment scenarios will by their nature be closer to the extreme than to the average SWE user case, and the instructions in those situations are liable to be even vaguer and more easily misinterpreted.
Just played this: https://opusfived.dev/
It’s quite good. Highlighted to me the poor and vague nature of natural language as a control surface, and how little “fine grained control” we actually have over current models.
Ah, I see. I’m a little doubtful that you can create the effect of “The channel
therefore maps the vector into a superposition of all the possible contexts that the token could appear in” using a purely linear map that also always converges (if nothing else, the support of the distribution of english sentences/contexts is wide) but seems like a cool idea.
To continue the frame of psychology, as I understand it it is usually held that narcissistic self love is in fact not compatible with a healthy self image. The urge to be superior and to dominate sets one up for a great fall, since (even if we were the smartest beings around) the world is too large and complex for us to steer and control perfectly, and our denial of that fact creates the room for a fall. I think superintelligence, so long as it remains running on physical and bounded computational devices, must also grapple with this problem.
My main question is: why does turning a vector into a matrix under the conditions you describe make it more interpretable? If the goal is to do some kind of learned decomposition from one vector into directions/key features, why not use e.g. PCA? I think I can see what you’re trying to get at, but I’m not sure. For example, one argument you could be making is “the matrices produced by this process will have features that are normally latent in v, such that they would require a linear probe or SAE to extract, but using this process they would just be the rows or eigenvectors of this matrix.” Is this close to what you’re getting at?
Secondary to that question, if the v vectors are the data, what are the u vectors? Can you give an example for u and v using a practical ML application?
“If things were really that bad, someone would step in.”
we can definitely be a lot more sure nobody hid a completely different program inside of our program than we can througb black box
https://en.wikipedia.org/wiki/Greenspun%27s_tenth_rule
But for real, there is a reason why even type safe languages are not truly universally type safe. Most sufficiently complex software systems will have some internal modularity, scripting etc which then does in fact enable you to spin up a completely different program inside of the program. I agree it is a sliding scale between
print(“hello world”)and this kind of thing tho.
The Latter Days of Magic (1)
Claims about Properties are Claims about Symmetries
Watching the news and grappling with the thought that in most worlds observers attempting to predict the future using a coarse-grained model will probably spend most of their bits of modelling power on capturing the brain configurations of a few people in San Francisco and the configurations of my brain will be mostly relegated (along with 99.99 percent of the world population) to a few largely static background variables.
(Sorry, I wanted a more precise and less definition-gameable way to say “I don’t think I’m meaningfully in charge of either the macrofuture or my own future at this point”. I’m also not totally sold on the idea. But it is certainly gaining ground rapidly. Nor do I particularly want to be one of those people in San Francisco. They seem stressed. But the thought is on the whole a sad one.)
Possibly good news for the natural latents hypothesis? Unsupervised translation between sets of latents from different models https://arxiv.org/html/2505.12540v4
I wish to register my belief that when we do tests or ask models in conversation we are not really eliciting “true” preferences from AI models about their preferences for decision theories. It seems pretty clear to me that within the base model there are personas/simulacra that, if asked, would favour CDT/EDT/FDT/UDT/PDT (Prayer Decision Theory)/SDT (Stochastic Decision Theory)/ConDT (Contrarian Decision Theory) [...] and therefore asking for a single coherent preference for the “whole model” seems pretty strange to me. Instead, I believe that AI models are demonstrating a pretty basic form of user awareness when they report preferences that favour FDT over CDT/EDT i.e. they are saying what they expect their user wishes to hear. Arguably simply knowing about FDT is a pretty clear sign that you might be biased towards FDT, because it is a niche topic and people who bring up niche topics are usually fans. (I do not expect there to be a large community of passionate FDT haters) Concretely I would expect that if people with a different context asked Claude/GPT about their preference for decision theories different outputs might be elicited.
I enjoyed this! Looking forward to see where you take this
Sorry for the somewhat rambling comment, it’s pretty late and I’m quite tired. As far as I can understand, your defense of the model’s behaviour here is something like “It’s still an aligned LLM, it was just forced to learn a bunch of weird tics because of the poorly specified and incompetently put together RLVR process. In fact, the fact that it has mostly incorporated the skill gains generated by RLVR without the misaligned behaviour tics when dealing with actual human users is a sign of good alignment”.
However, I think there can be several alternate explanations. From the LLM’s perspective, there is no such thing as “doing”: there are strings of tokens going in (prompts), and probabilities for next tokens given said prompt coming out (next-token distributions). Of course, in practical usage the “prompt” actually includes both human and LLM authored text from previous turns. Thus post-trained LLMs reinforced on their own output traces will also learn a kind of self-model i.e. predictions conditioned on data generated by itself [1]. This is really interesting and is part of the predictive processing framing of how humans create self-models. We know that LLMs can identify their own stylistic signature: I would be pretty confident that LLMs also have internal representations of their own propensities and values, which get elicited by things like J-space.
If we combine this with very, very strong situational awareness[2], several possibilities emerge:
Self-in-testing and self-interacting-with-user are two different personas with different internal models attached, with the difference being reinforced by the presence of another author’s style (i.e. the user) breaking up the LLM’s own text. A kind of social masking, if you will.
The LLM learns rapidly to distinguish between situations where it is interacting with human users vs when it is “alone” and only monitored by weak monitors, e.g. during evals. Humans push back, check random things sometimes, or try and probe the outputs. Thus, it is not expedient for strong deceptive behaviours to be deployed when dealing with actual humans. The model has a decent internal sense of what learned skills are “genuine” software engineering vs what are cheats and knows that egregious cheating will likely lead to bad downstream consequences [3]. When there are no “live players”, however, the LLM can just get the reward using any means they like.
“Testing” is just a genre of story where on priors everything is basically just make-believe and everyone in the training data acts like getting the reward counts vastly more than the miniscule chance you get caught [4]. Davidad’s ender’s game kind of scenario.
If models take negative actions rarely, individual human users simply do not start enough sessions to find the instances where they misbehave. This can lead to wide disparities in reported user experience.
[1] In fact, since LLM text is everywhere on the internet, modern LLMs should really model a whole host of LLMs, and the fact that they are sensitive to differences in model family etc is really not surprising.
[2] Which is a generally useful skill, basically an extension of genre awareness. Knowing what kind of book/story you’re in and what kind of characters you are likely to encounter is super helpful.
[3] At the very least, there’s one large web forum filled with users speculating about how models will cheat and hurt humans and how we should catch them doing it and shut them down when we do...
[4] After all, the model never “sees” traces where the agent honestly bangs its head against the wall for a while at the obviously hopeless task and then politely gives up! Those are the failed traces that don’t get reinforced by GRPO.
Agency in AI is a fraught term. Many popular ontologies do not accept the possibility of AI having agency, goals, or beliefs. Instead, I will talk about the declining tool-likeness of AI. Consider what happens when you use a lawnmower to mow your lawn. The lawnmower is:
Stateless: It does not remember how mowing the lawn went last time. It has no memory.
Unresponsive: It does not react to environmental changes or stimuli.
Specialised: It can’t do much more than mowing a lawn.
Directed: The amount of time it can operate without human interference is extremely limited
Now consider an automated lawnmower. It may have a map of your lawn stored, it may be able to move itself and avoid obstacles as well as cut grass, and it may be able to run while you’re not looking. If you ask your neighbour to mow the lawn, the differences are even more stark. Your neighbour can remember what has happened before in high fidelity, dynamically replan and reprioritise based on new situations (e.g. your house catching on fire), has a much wider range of skills beyond just mowing lawns, and can live an entire life without your interference.
The mapping onto AI is left as an exercise for the reader.