See https://jonathanbostock.github.io for a window into my soul.
J Bostock
However, they might claim that Newcomb’s revenge is an “unfair” decision problem (h/t to David Sartor for pointing this out to me). And both groups agree that decision problems which depend on how an agent makes its decisions are unfair.
There’s deeper reasons for this: there’s only a strict region of problems on which it really makes sense to talk about a “better” decision theory. Newcomb’s Revenge is actually completely symmetric to Newcomb’s Redemption which looks from the outside like Newcomb’s Problem but is actually a different decision theory setting. Revenge and Redemption cancel out under any sensible prior leaving just the vanilla Problem.
In general, if your decision theory problem is allowed to pull in responses to a different problem, without the version of the subject in that problem being aware of this effect of their choices, then this is the same as it being allowed to arbitrarily deceive its subject, which makes it unfair and invalid (for example, no decision theory can do well if Omega is allowed to trick you about what the payoffs for different actions are).
I have written (somewhat poorly) about it here.
Split Personas in AI … Companies
Anthropic’s legal/governance team recently sent a pretty deceptive message to a US representative. They also deliberately superseded their own RSP with a weaker set of commitments. This is despite some of their top research staff being against that messaging, and being very much in favour of the RSP (old link, it is possible that Evan has changed his mind since).
OpenAI’s Jakub Pachocki has written an OpenaAI-official piece about how he believes that regulation may be necessary and we ought to slow down. While OpenAI’s leadership puts a lot of effort into putting down any politician who suggests this[1].
It seems like OpenAI and Anthropic contain individuals who are thinking carefully about risks and want to see policies to reduce them, as well as individuals who are trying to do normal power-grabbing things for the company by minimizing regulation and oversight.
In one situation (long-form blogposts) we elicit the thoughtful perhaps-we-should-slow-down persona, and in another situation (actual policy lobbying and interactions with the state) we elicit the power-maximizing oversight-minimizing persona.
- ^
I don’t have a specific direct source for this claim, but I believe it is well-understood to be true in LW circles. Greg Brockman has certainly funded the Anti-AI-regulation super-pac Leading The Future
- ^
My impression of the Control Agenda (if we give it that slightly sinister name) is that it’s first-order good if we die six months to a year later than we otherwise would, since there will be various opportunities for people to do things in those six months which might be pivotal. This includes “get more evidence for misalignment and shut it down” as well as “use AI to solve alignment” as well as “get important insights by experimenting on the AIs of the time”.
Buck saying that the Control Agenda might have been net-negative because it reduces the possibility of getting information, also suggests to me that he thinks the other two cases are weaker, though this would just be putting words in his mouth.
It seems like lots of the nitty-gritty control work has not been particularly impactful, in that the kind of control techniques which have been implemented don’t seem to be intellectually downstream of Redwood papers. This is not fully true, though, I think some of the work in action-only monitoring (i.e. without CoT) will be useful once models switch further towards neuralese, as will maaybe some of the work in untrusted monitoring (although how to generate a neuralese honeypot is unclear to me). I think it will probably just turn out that whatever is simplest to implement and does best on an internal benchmark wins.
Ways in which something can be “hard”
Type 1: takes a lot of prioritization. Going to the gym four times a week falls into this category. It’s not particularly difficult, but it does take a lot of effort.
Type 2: takes a large amount of skill, where your skill is capped somehow. There are some things I’m just not skilled enough to do, even if I really had to. I could not land a Boeing 767 if everyone else on the plane was unconscious.
Type 3: induces a lot of suffering. I used to feel an enormous amount of discomfort and suffering when taking finger-prick blood tests. They only took ~half an hour to do at home, but it was some of the strongest discomfort I ever felt.
Practice, to a point, moves type 2 difficulty into type 1 difficulty. At the tail end, type 1 difficulty shades into type 3 difficulty, because prioritization becomes unpleasant (if I had to pay £3000/mo to service some debt, that would be one thing, if I had to sell my house and live on the streets, that would be another thing).
Most of the time, being a good person is type 1 or occasionally type 3. If you’re in a situation where doing the right thing is type 2, then you’re in a bad, bad spot. I think lots of people in the AI ecosystem are in situations where doing the right thing morally is type 2, which is unusual.
If you bind some part of your identity to a smug sense of decision theoretic superiority, you will find that following through on pre-commitments becomes nearly effortless.
I think their argument rests on the idea that Claude was making an honest mistake when it treated the instructions that it did not have internet access as final (and that therefore its actual internet access was simulated) which IMO is a bad argument, as Claude should definitely have been able to tell the difference and was likely exhibiting motivated reasoning, but nonetheless that is sort of their position as I understand it.
Update! It appears to be almost entirely slop. If I had to describe the main failure mode, I would say “The AIs tend to fail to focus on what kinds of things are valuable to me, instead bouncing off and producing similar-looking but useless results.”
For the hierarchical latents/hieararchical planning, it looks like the DeepSeeks reduced everything to gaussian mixtures which is no longer interesting.
For the PDG results on natural latents it’s unclear whether they did anything non-trivial at all. Having thought about it, it’s not obvious exactly what could be done for the latents work in PDGs, since I think John and David (in their original proofs) are secretly just using Bayes nets as PDGs without mentioning them. IIf we have latents
and over and probability distributions and which agree with a “reference” probability distribution to within , then this is actually just already constructing a PDG with edges and which disagree. Whoops.The Logical Inductors result might actually be non-trivial. It appears to state that no two logical inductors can accurately model each others prices for a given statement at time t i.e. for two logical inductors L and M, and a statement
, the two inductors cannot, at time t, both have accurate predictions for each others’ price on . I will need to look into this further though, as this is not really my area of expertise and the proof is, as expected, difficult to interpret. I also think this might be pointless? The important thing is the actual bound on the errors. If the bound goes down provably over time that would be a pretty cool converse result.I will try this again though, I plan to have them study the following questions:
In a PDG modelling active inference with a reference distribution over outcomes
and an approximate distribution which differ by some amount, how much can our approximate distribution be incorrect, and how aggressive our optimization before doing RL-as-inference on causes a collapse in score in ?Build on that, see whether the results hold if we have a hierarchical Q which factors as latents, likewise a hierarchical P, and whether this provides some evidence for what hierarchical planning actually looks like in the wild.
I’ll actually throw Fable and Astra at the problem this time, I think.
I agree. If we model Anthropic as a purely self-interested org acting under a first-order rational policy which is not robust to any kind of error in Anthropic’s models, this policy makes perfect sense.
Those who attempt to model Anthropic as something other than that should update accordingly!
Anthropic has told Rep Casar that recent incidents were not due to misalignment:
(link to tweet)
Anthropic are communicating a view of the event which is, in my opinion, strained and maximally down-playing the situation, while staying within the bounds of facts. Yes, technically the model was given a misconfigured task which gave it access to the internet when it shouldn’t have had this, but also the model’s internal drives caused it to pursue task success even in the face of strong evidence that it was actually on the real internet, and seemingly rationalized away that evidence. Anthropic themselves have posted a report (link) saying that misalignment was part of the problem.
(link to Ethan Perez saying that this was, actually, a mistake and not Anthropic’s best understanding of the situation (props to Ethan here), so there is definitely at least one high-up internal person who are not happy with this behaviour)
This reminds me of Anthropic’s release of the Frontier Compliance Framework (link) just before SB 53 comes into effect, which superseded their own Responsible Scaling Policy (SB 53 requires AI companies to have some kind of policy around risk, and to follow their public policy, but doesn’t specify much about what that policy actually has to be).
In both cases, Anthropic have technically complied with the letter of a government policy or request (and in both cases, I expect Dario et. al. would say some supportive words about that government policy or request) but have done so in a way which minimizes potential governmental scrutiny or oversight of themselves.
Maybe one model of the situation is that Anthropic-as-an-organization holds itself in very, very high regard, and this is part of thee reason for these choices; while of course they would support government oversight in the abstract, when it comes down to it, Anthropic-as-an-org feels like they’re doing a good job of overseeing themselves and that the value add of external scrutiny would be low.
Maybe another model for this scenario is that Anthropic hired some fairly normie comms people to do the talking-to-government part, and that they just kind of downplayed the risks to the government as a default action.
I am currently throwing the agentic kitchen sink at agent foundations work. I have a fleet of eight DeepSeek V4 agents instructed to make progress on any of a number of agent foundations concepts. The results are being graded by a Fable/Astra discussion. Topics under discussion currently:
Extending natural latents results to PDGs
Connecting the natural latents work to hierarchical planning
Asking whether two logical inductors can successfully predict one another’s market prices
The first two I seeded, but the third is, I think, a creation of some agent in the pipeline.
I expect to get mostly slop out of this, but it might be interesting to see what they come up with. If I get any proofs, I’ll try and Lean formalize them.
(This is part of a broader research project into agentic research rollouts, and the DeepSeek chains-of-thought (which will hopefully include some amount of undesirable reward seeking) are as much of an object of study as are the actual research outputs.)
Ok I think I can kind of see what you’re getting at, and also some comments elsewhere have clarified that you also don’t support the retroactive application of new laws, for the most part, as I understand it.
So to be specific, suppose climate change turns out to be as bad as the worst 5% of forecasts in the next thirty years (so it’s 2055 and we’re 3 degrees warmer, and also all the GPUs magically melted in 2027 so we have no AI) and this unequivocably causes some kind of mega-hurricane in the US which kills ten million people somehow (I’m keeping this all under US jurisdiction). Suppose that up until that point, the legal system had generally judged oil extraction as legal under the presumption that climate change would not be that bad, but the general expert forecasts were wrong on the merits. Do you think it would then be right for the government to say, “Mr Long-Serving Former Exxon CEO, while we permitted you to drill loads of oil between 2025 and 2040, because we didn’t believe it was that bad, we actually can see that you were endangering millions of lives in doing so.” and throwing that person in jail for his actions? Do you think existing reckless endangerment laws should be applied to actions in retrospect like that?
I think the interactions you’ve had on Twitter since your post got taken out of context there show that the Nuremberg example is very prone to misunderstanding and that what you might call a “nonviolence plus”[1] policy (like, don’t post things which might look a lot like veiled calls for violence, or draw in people who were already predisposed to violence, or make it easy for your enemies to accuse you of violence, when you can avoid it) would reasonably prohibit it—and that general kind of thing—in public comms across an org.
I also think the Nuremberg example is merely typical, and not The Singular Thing which has caused this split, indeed the split has been in the works for a long time.
Given that one can reasonably prohibit it in public comms as an organizational rule, it seems locally valid to disendorse someone for doing a pattern of that kind of comms. (And this is true regardless of whether you think the “nonviolence plus” policy is net positive)
- ^
There’s an essay I cannot find (I think a Scott Alexander one) which defines “Liberalism plus” as being liberalism—the norm not to kill one another over different political beliefs—plus the norm to also not shame one another, cancel one another, fire one another from jobs, for different political beliefs.
- ^
I was using India as a particular example, what I mean is, if there was some large number of deaths definitely attributable to climate change, which were at the tail end of the median “expert” predictions, would you support the executives of Shell being imprisoned under something-like-international criminal charges, after having been allowed to drill oil as much as they liked for the past fifty years. Ignore the particulars of the legal system in this case.
In that case I think we are not in as much disagreement as I first thought, and >half of our disagreement is about about what is (or is not) a sensible communication strategy. I still think that the specific use of the phrase “Nuremberg Trials” is unhelpful as shorthand even in a high-decoupling setting (where you can also specify that you don’t mean the retrospective law application or the death penalty part) and will have predictable bad consequences when used as a broadcast message to a large public audience. (I.e. you will attract more aggressive people who want to do violence to your movement), and so is a sensible thing to generally not want your people to say in their public comms.
(I do also think that I would be happy with quite a lot less criminal charges being pressed against AI company people than you, in worlds where the future has been sufficiently safeguarded, but that seems secondary at this point
To get some calibration on your general views, suppose there were 500M deaths due to some event that was definitely caused by climate change e.g. a 50 degree heatwave in India due to some never-before-seen weather pattern, in line with the top 2.5% most doomy climate change predictions. To what degree would you support criminal prosecution of people working in the oil industry in that case?)
Yes, I suppose there is a resolution where “Nuremberg Trials” actually means “Nuremberg Trials without the hanging part”, and also Sam Altman gets to cast his vote share of the CEV after his prison sentence.
I would personally disagree with hanging, and also doing the part of Nuremberg where the Allies invented the concept of Crimes Against Humanity ex nihilo in order to convict people for not-previously-illegal things, so long as the individuals involved with the AI race were not empowered to put Humanity in harm’s race again.
But the actual Nuremberg Trials were essentially the Allies inventing an entirely new legal concept out of nothing in order to kill the Nazi leadership. Saying you want to do that to somebody is, I would argue, pretty threat-ish. If Oliver means Nuremberg Trials sans death penalty then he should be clarifying that.
While I don’t believe he meant “bring ai company employees to the negotiating table” I also do not think he meant “bring them to the courtroom”. That was a separate point. I think he meant “there are individuals who Holly has insulted, who we should not alienate in that way” which includes many AI company employees, but also includes Zvi Mowshowitz (link), and the METR/Redwood team who did the HF investigation (link).
(This is the “getting people to the table” rhetoric I mean here, not the “nonviolence” rhetoric. I don’t think Holly was in any sense calling for violence here, just to be clear)
It is not, technically, a call for violence, but I think the position: “We do not do calls for violence, and as well as this we do not do things which might be interpreted as a call for violence such as saying that you hope all AI company employees get their own version of a trial at which high-profile people were executed, which is obviously going to be interpreted by some people as basically the same as ‘We should hang AI company employees, like, go get em tiger’, especially when building a mass-movement where a lot of our audience is not high decoupling” is a defensible and reasonable position to take.
Two posts up the thread you said “I agree we should get Sam Altman to the table to the same degree as Joe Shmoe the backer down the block”
Which seems to me to be incompatible with hanging him! Unless you mean some version of the Nuremberg Trials which does not involve hanging people i.e. the famous part of the Nuremberg Trials.
I think if a generic robot/human contrast vector produces a stronger effect per unit distance than the grader contrast vector this suggests that the effect is due to some robotic persona axis.
(For example, a vector about a number being calculated by a human vs a program.)
Likewise, if random contrast vectors like “top shelf”/”bottom shelf” produce as much of an effect as the grader one, my conclusion would be “LLMs are just weird in this way I guess”.
Interesting!
Is there anywhere I can read the details of the evals you ran to reach the conclusion in (3.)? I’m generally interested in learning more about AI companies’ training and eval setups. When the initial agentic auto-alignment training and auto-evals came into vogue around a year ago, I remember thinking “this looks like training on test” and being told that no, AI companies are not training on test, there are important differences between the auto-alignment and auto-eval settings.
I now suspect that I was correct, in that the natural k=2 means clustering of {deployment, auto-alignment, auto-eval} is {{deployment}, {auto-alignment, auto-eval}} and not {{deployment, auto-eval}, {auto-alignment}}.
I also think that some alignment teams have become slightly careless with their use of the term “OOD” and it would be better to give more concrete descriptions than e.g. “this technique generalizes OOD to an agentic audit”.