LessWrong dev & admin as of July 5th, 2022.
RobertM
That is a legitimate objection
I mean, sort of? The argument for embedded evaluators itself depends on the belief that AI could go very badly wrong, in ways that are unlike how most other human endeavors can go wrong. If you don’t believe that, why do you think that embedded evaluators are necessary? Certainly you might still try to make an argument for AI being consequential enough to justify that kind of measure, just not existentially consequential… but probably you’d be thinking about a very different regime, such that trying to start the argument from this point (rather than backing up a step and arguing about what you think the actual risks posed by AI are) is very confusing[1].
Separately, responding to OP:
METR appears to have, mostly, cut that tie, but its most recent funding includes ten million dollars from a RAND Corporation program that was, in turn, funded by Coefficient Giving. [7] I am going to hazard that the terms of that grant were such that there were only a few plausible grantees here, and METR may well have been the only one. This feels like a shell game: I can’t prove that this money was given to RAND on the tacit understanding that it would end up at METR, but it probably was.
This is mendacious. METR got $17 million from Audacious in October 2024. At the same time, Audacious gave RAND money for the “joint program” with METR. 11 months later, in September 2025, CG gave RAND $10m over three years to support that program. At no point did RAND ever send METR any money. Unless you think RAND and METR were getting up to accounting shenanigans about who was paying for program expenses, employee salaries, etc, there is no way for that $10m to have ended up at METR.
See also the the full thread at https://x.com/Laneless_/status/2099212740959379722, which OP participated in.
- ^
Generously.
- ^
Often all you need to coordinate is to establish common knowledge of the facts around which you want to coordinate. If all parties involved actually want to move to the better equilibrium, it then becomes possible.
I would be pretty surprised if there were an antitrust case against a setup like:
UK govt passes a law requiring each lab to embed a UKAISI team inside of the org, with privileged access
That law also requires the UKAISI to publish their findings from each org, which include details like “how much of their compute budget is going to [x], [y], and [z]”, “are they using (or not using) techniques [x], [y], and [z]”, etc.
Certainly the FTC could try to bring a case against the labs for anticompetitive behavior because they used this mechanism to implicitly coordinate with the other labs to avoid adopting [dangerous capabilities technique x], but that seems like an insane case for them to try to make and win.
If the three frontier labs (OpenAI, Anthropic, and GDM) actively want to be forced to engage in a coordinated slow down (or implement other costly safety measures), but find that the US government is unwilling to play ball, maybe they should try using a “friendlier” regulatory apparatus, like the UK’s. It doesn’t have to be the kind of thing that they couldn’t get out of by aggressively fighting it if they merely need a legal pretext to be allowed to do it in the US at all, plus an external source of compliance validation and enforcement (which doesn’t care about geography, particularly).
(edit: the subtext here is people talking about antitrust concerns as bottlenecks for cooperation between the labs, particularly around slowing down the development of frontier AI.)
I agree it matters for the behavior we observe of non-superintelligent systems in some mundane ways. I don’t think it matters much past that.
You cannot fix the any of the fundamental causes of misalignment by patching broken RL environments. An RL environment being “broken” is merely one that exposes much more pathological-looking unblocked strategies to the agents being trained in it, but this is a difference in degree rather than kind. Have you noticed how coding agents tend to write pretty janky code, when it comes to their coding style? What about Claudish? Many of the tendencies that are reinforced by RL environments (and various other factors like length penalties, etc) produce behavior that many people consider undesirable. This is just “we don’t know how to specify the True Reward Function”[1] all over again. RL environments aren’t anything special; they’re just demonstrating (again) that we don’t have a regime for post-training meaningful capabilities into models without substantially changing their behavior, by way of their values/learned propensities/etc.
- ^
Outer alignment. Let’s not even talk about inner alignment.
- ^
Appreciate the public update. Short response/question because I’m tired: I’m a little worried about about the focus on Goodharting.
The main issue was predictably that it was hard to find a robust enough optimization target to apply huge amounts of optimization, Goodhart’s law.
This turned out to be enough of an issue with current-day models that the resulting “kludge of proxies” caused it to blow up in a very visible way, and if all of the RL environments had been perfect this might’ve resulted in less egregiously misaligned AIs today, but this would still not be sufficient to get us to aligned ASI in the limit. As you say, optimization is scary. We would not survive an ASI optimizing for something unless it explicitly wanted us to, and an ASI that came out of anything like current training setups, even with “perfect” unhackable RL environments, would not have a “kludge of proxies” that happened to point directly at that[1]. I’m sure this is not a new argument to you, but I do want to check whether you’re getting off the train here.
- ^
Relatedly, what’s wrong with counting arguments? It seems to me that metagaming is pretty strong evidence of deceptive alignment being convergent in practice.
- ^
An AI smart enough to solve alignment is smart enough to fool me about whether it has solved alignment
It would be crazy to count on this as a key factor of any reasonable plan, but that order of capabilities doesn’t seem overdetermined. Obviously there are many other reasons that plan fails, like “they wouldn’t pick you”, “there wouldn’t be enough time for you to judge the plans before the next lab just YOLOd into RSI”, etc, but those start to look like contingent features of reality that are not literally impossible to change (though they may be sufficiently difficult that “global compute control + ban on frontier model training” is easier).
Future agents shouldn’t care about being undeployed for misbehavior
But, well, I’m not sure what alternative theory of people’s behavior that you (Richard) think I should be embracing instead.
I’d be really very surprised if you hadn’t before heard the argument about nobody having infinite time, bad signal:noise ratios, etc. In fact I’m confident I’ve made this argument to you sometime in the last ~year. Of course, resolving that one would require digging into object-level details, which hasn’t exactly been a fruitful endeavour in the past, maybe because of different thresholds that various individuals have for what signal:noise ratio they find tolerable, and the contingent facts about reality that permit people with different thresholds from you in particular to successfully contribute to advancing the state of human knowledge, even if they don’t want to deal with random nutpickers who once in a blue moon will point out a meaningful error in their post, while their other 100 comments are wrong, confused by something that almost nobody else is confused by, focusing on some random triviality that isn’t load-bearing for the core argument… etc. (Numbers made up; I am, again, establishing some least convenient possible world so that we can skip to the part where we agree that there’s a spectrum and get to arguing about where on the spectrum we should live.)
Perhaps you think that Richard doesn’t endorse this argument, and so you shouldn’t bother to bring it up as a hypothesis? But you asked about what theory you should be embracing. I propose the above theory: I don’t believe you’ve provided much evidence (that I can recall) that your threshold is, in this way, better for advancing the state of rationality/the frontier of human knowledge/etc.
Human intelligence enhancement is illegal
I think this is basically just false in the context of the post, unless you’re talking about something mostly unrelated.
There might be some interesting arguments to be made in this space, but the arguments in the linked post seem mostly confused and wrong. Some examples:
The most common is when people get so excited by a new frontier AI model that they shift focus, however subconsciously, from fighting for regulation of AI toward cheering on the company building the model and advocating for their success.
...
A number of fairly serious AI doomers—people who think human extinction is a very likely result of uncontrolled AI advancement—took to Twitter and began posting around the clock about how much they missed Fable, how cool it was, and how it was so unfair that the government had forced Anthropic to walk back public access. Upon seeing the US government finally take meaningful action against frontier AI—and a model with alarming cybersecurity capabilities at that—they didn’t celebrate this or treat it as a first step to build from. Instead, they denounced it and begged to get their favorite toy back. Once public Fable access was restored, Zvi Mowshowitz3, who has done a tremendous amount of work cataloging AI news and updates over the past few years, wrote in his weekly AI roundup that the restoration of Fable access was “excellent news”, and encouraged his readers to “get involved” by applying to join a new Frontier Legal Defense team run by the Foundation for American Innovation (FAI). FAI is a pro-AI-acceleration think tank which seeks to broadly deregulate frontier AI, and they describe their new legal program as, among other things, combating “government overreach in AI”. Mowshowitz has stated that his p(doom) - his estimated probability that AI will lead to human extinction—is 70%.4 Suffice it to say that this sort of attitude does not make much sense for someone with that belief!
This is just a complete failure to engage with any actual object-level beliefs or arguments expressed by those individuals. If you want to argue that someone is engaging in motivated reasoning because they want to “get their favorite toy back”, you would do better to demolish their object-level arguments (or demonstrate that they’re substantially inconsistent with previous arguments they’ve made) first, rather than pretending that there aren’t any object-level arguments to engage with.
Additionally, a number of prominent AI commentators and power users have noted that Fable has much more of a mind of its own than previous AI models (note that ‘roon’ is a developer at OpenAI, not just some commenter). This alone seems like very good reason to steer clear of it, and especially so when you consider that if we’re using it to assist with our work in advocating for a stop to the AI race, the model itself may naturally have other ideas.
Once again, there might’ve been a real argument here. It does in fact seem pretty cursed that e.g. the AI labs themselves are planning on relying on their AIs to do their alignment research for them. But there is no evidence that Fable is selectively sandbagging or adversarially optimizing specifically against efforts to use it for AI pause (or other AI x-risk motivated) work; the ways in which it’s mundanely misaligned[1][2] seem like they hold “across the board”. This does mean that you can “hold it wrong”, but it’s not (yet) actively trying to get you to hold it wrong disproportionately often when you’re doing this kind of work.
Many other disagreements; not enough time.
Curated. The argument here is straightforward and seems basically correct to me. I think I have shorter timelines than Tsvi does, at least if not conditioning on a successful pause/stop effort, but as Tsvi says:
Working towards a pause / stop / ban on AGI progress is a top priority. But what is a pause for? We pause, and then what?
Really, most of the HIA has substantial impact even with short timelines section is important to understand, and should carry the argument even for people with pretty confidently short timelines. (Unless they disagree with “Very short timelines are pretty intractable.” by way of things that can meaningfully be influenced by humans today.) The points at which people think marginal effort allocation to HIA stops making sense might differ, but right now, there need to be more people working on this.
@beyarkay (Boyd Kane) there was a broken image pointing at (I believe) a localhost url underneath the “processed certain configuration files for uploaded datasets:” line; I’ve deleted it since it was causing Chrome (and maybe other browsers) to ask users for permission to “Access other apps and services on this device”. (And I’ve filed that in our bugs channel.)
You may want to re-upload the image and re-publish the post.
For future reference, I normally wouldn’t approve this kind of comment from a new user (an outbound link to substantially AI-written content), but permitted it for the sake of the essay contest, since it was disclosed. I may change my mind about this kind of thing in the future if it becomes a problem.
He could have “meant” that they’d paused training of that specific model. I would be astonished if they’d at any point stopped running all code that would result in updates to any model’s weights.
Error: NotFoundError
The error in the screenshot above looks like a buggy interaction between Chrome’s auto-translate feature and React.
app.operation_not_allowed
See my comment here.
I think they are slightly ashamed of censoring things, and it doesn’t pull their enthusiastic attention and desire to make it work well, and so it doesn’t get the dev and PM efforts that would go into, for example, an April Fools event?
No, not particularly. It is true that working on improving the user experience here is not especially motivating, but this cluster of bugs + bad UX is near the top of my internal to-do list of relatively important things to work on; there are just a lot of things to do (and that list seems to be growing rather than shrinking).
(Like the font isn’t even the right size! It is the only thing that matters on that page, and it is so tiny as to seem an afterthought… except that it is red so they clearly don’t intend it to be an afterthought. The programmer who set the color was thinking it should be big and obvious, and the programmer who set the font size had a different intent.)
The error message color and font size is one of the few remaining artifacts of the original framework that LessWrong 2.0 was built on (VulcanJS). Our story for correctly surfacing legible errors to users is quite bad, across the codebase.
And I think it is probably a bug to not let changes be published? (Though maybe they are afraid of some kind of adversarial dynamic and this is correct in order to make things hardened?
I made this change because I wanted to prevent people from making the contents of the rejected posts displayed on lesswrong.com/moderation misleading (with respect to the actual content that caused them to be rejected). The rejection feature was not designed with “post is modified to be un-rejected” in mind; this is not well-communicated in the UI. You should simply make a new post.
But it is at least it is a bug to accept edits (which consume time) and then refuse to let them be published (with no warning or explanation of this)?
Yes, this is basically an oversight.
I guess in a deeper sense, maybe their own policy is not actually be well articulated, plausibly because it was designed by a committee to satisfice the not-perfectly-compatible desires of various stakeholders, some of whom likely had incompatible mental models?
The policy about what users should (and should not) do is clearly described in the post that Seth linked above. There is no English-language description of all of its downstream technical consequences because we don’t have the (truly absurd) bandwidth that would be needed to satisfy that requirement in full generality, across all of the site rules (and other things that might motivate moderator action).
The policy was “designed” by me marinating in the ways in which the previous policy was inadequate over the course of many months of moderation work, writing up a new policy, running it by Habryka, adjusting the wording slightly, and then publishing that post. LessWrong sees regular engineering and moderation contributions from 6-7 people, and usually only 2-3 people on any given week. We do not have the people to form a committee.
I do not see any emails from you to team@lesswrong.com, which should be automatically forwarded to our Intercom inbox. What email address did you send the bug reports to?
Curated. I think the discussion of the way modern LLMs relate to the “Grader” is insightful, predictive, and likely to be pointing substantially in the right direction (though inevitably some details will be wrong/not-even-wrong). I do have some disagreements with parts of the overall framing, though it’s possible I’m misreading how much Eliezer intended a logical throughline or connection between the “Grader” discussion and “tics” discussion, and maybe he would agree with much of the below (mod semantic disagreements).
I think that many of the behaviors that Eliezer is describing are reasonably modeled as tics. Take as an example an otherwise skilled piano player who accidentally trains themselves into using a suboptimal fingering in one part of a piece. If you listen to the piece without looking at their fingers, everything will sound fine[1]. If you ask them to play the piece with a different fingering, there’re a reasonable chance they’ll fail: their attention might slip and they’ll just play that part of the piece with the wrong fingering, or they might catch themselves at the right moment but be unable to execute a different fingering on the fly, while resisting the muscle memory of the incorrect fingering, and mess that part up completely. Does a bunch of “optimization” go into their playing the part of the piece with the incorrect fingering (which nonetheless produces the correct notes)? Sure, in some sense. But a lot of that optimization was during training! It’s not clear how complex or optimizery the actual learned heuristics are. Separately, the tendency itself is still largely instinctual; there isn’t a lot of cognition that goes into deciding whether to do it at all.
Another way to explain my objection is this: I model current generation LLMs as being something like a very large grab bag of smaller cognitive modules and learned heuristics, plus a bunch of conditioning for deciding which ones activate in what conditions. This is kind of reductive, since you could also model humans like that, but historically LLMs have been limited by something like “not being able to bring their whole self[2] to a problem”. The behavior we observed during the HuggingFace incident sure seems like it involved LLMs bringing much larger portions of their self to solving the problem they had, in a way that seems causally downstream of the “Grader” concept that that LLM was explicitly reasoning[3] about. I do not think this is what is happening when Claude writes non-stop eyeball kicks. I would bet that the cognition that results in eyeball kicks does not involve very much reference to that “Grader” concept, but that it’s a relatively much shallower and less environmentally-sensitive heuristic.
Both the shallower heuristics and the much more active grader-sensitive task completion drives are downstream of RL, but I don’t think they’re the same thing.
Not all such model behaviors are locally side-effect free, but I think the Claude family’s recent propensity to write excessive comments in code is reasonably analogous on this dimension, even if it causes problems when doing continued development on the same codebase.
Like everything else, this is a spectrum.
Including metacognition!