Previously “Lanrian” on here. Research analyst at Redwood Research. Views are my own.
Feel free to DM me, email me at [my last name].[my first name]@gmail.com or send something anonymously to https://www.admonymous.co/lukas-finnveden
Previously “Lanrian” on here. Research analyst at Redwood Research. Views are my own.
Feel free to DM me, email me at [my last name].[my first name]@gmail.com or send something anonymously to https://www.admonymous.co/lukas-finnveden
If Joe still has Anthropic equity, then that’s a listed example of a tier 3 COI so would have to be disclosed in any Anthropic investigation he’s involved in (and also non-Anthropic investigation unless a non-conflicted employee takes responsibility for his work).
(Ex-employee without equity isn’t listed as an example so would fall back to the more general principles discussed, especially “Write down the full details of any potential conflict(s) of staff working on the risk assessment, and how those staff contributed. If we published that in full, would a reasonable person not affiliated with METR have materially less trust in the assessment?”)
If you want to feel the alignment difficulty that the most hardcore “doomers” feel, you only have to consider that they think the checklist looks like this:
Solve moral philosophy, finally & forever
Impart the correct moral philosophy to the AIs
(These days, it’s common to insert a step 0: build a very-capable-but-not-quite-too-dangerous AI, to then achieve step 1 for you.)
Any hardcore doomers who can comment on whether this is what they think the checklist looks like?
Re last paragraph: But would it be substantially more compelling for Chinese companies?
Ie, I totally believe that the US can build much stronger regulations for its own companies. But David’s post was about international governance regimes, not domestic ones.
I don’t know, I’m skeptical that this global standards body, regulating both China and the US, will have enough power to fully stop development when a company’s evaluations or internal cybersecurity is deemed to be not quite up to their standards.
What about when they’re egregiously not up to the standards, or equivalently, not up to pretty loose standards?
Ie., a country’s scientists comes back to their home government and says: This agreement is being clearly violated. I think we’d have a scientific consensus saying so in our own country if only we had enough time to establish that. The other country is putting us in serious danger. Please ask them to stop or to put in more effort into showing we’re wrong or else.
Do you still think a response greater than 3% of the company’s revenue is unlikely? Or that it’s just not very valuable to be able to act on this kind of thing because for every regime you can realistically set up, you’ll be able to do _whatever you want_ with plausible deniability because there’s so many places you can fudge numbers and pick favorable evals? (Such that there’s no value add from being able to rule out clear stuff.)
Also, I think a plausible reaction to even mild violations could be withdrawing from the agreement, which can be quite bad if both countries benefitted from the agreement being in place.
because I read them as almost universally “Eliezer vs. EA” conflict, not “rationalists vs. EA” conflict. He expresses frustration at the rationalists to a substantial degree as well, and my inner Eliezer would not say things like “rationalists vs. EAs”.
I’m not sure why you think this. For example:
Wei Dai: “you’ve repeatedly drawn a line between ‘my people’ (which I interpret as MIRI/rationalists/LWers/people who have absorbed the Sequences) and EAs, then criticizing EAs as a group” … Eliezer does not contradict him (so presumably grants the premise?) and instead elaborates: “Carl Shulman and Luke Mueulhauser were very Modesty-Argument / ‘outside view’ people, which tightly corresponds to getting sucked in by the EA cluster once it exists.” … “Carl was a rationalist before there was an EA/rationalist dichotomy. After the EAs showed up and started being Modest, it started being clear after a few years that Carl was Modesty-seduceable and hence leaning increasingly OpenPhil over MIRI.”
Or Holly: “There’s this common EA/rat argument that it’s important to get good ppl in ‘the room’ so they can intervene at a crucial moment.” Eliezer replying “I’d call it an EA argument. You don’t see me desperately scrabbling to be in the room at any cost in harm.”
Also, re expressing frustration with rationalists: making the corresponding twitter/LW searches for “rationalists” instead of “EA” I find no expressions of frustration. Indeed, on twitter, most tweets are ~defending rationalists from other people sneering at them. (Plus one tweet saying he doesn’t usually self-describe as a rationalist, though his other tweets seems to mostly accept that other ppl will associate him with the label and therefore use it in a way that includes him eg the Holly tweet response.)
(Edit: Anyway there’s a lot of stuff happening in the world right now so I should probably pre-commit to not spend any more time responding to this thread over the next week at least. You’re welcome to respond or not to this last message as you see fit of course.)
Your first two links are wrong.
Thanks for this. Will think more about this, but one quick question:
Labs also report CoT uplift specifically, defined as the ratio of (covert sabotage success rate with CoT visible to the monitor)/(covert sabotage success rate with CoT invisible to the monitor). [...]
I think CoT uplift is also potentially a thing we could make a core part of what we’re trying to preserve in the Monitorability Report—it’s a bit further away from the core risk-relevant thing (at least compared to the main trendline thing above), but I think it still seems very important and there are a lot of diffuse benefits that labs get from just having CoTs everywhere for people to use (e.g. for monitoring, observability, interpretability, etc.), so preserving that I think would be super valuable if possible.
Since these diffuse benefits aren’t specifically about sabotage anymore, I would have thought you wanted to track them using something like (with-CoT capabilities across a broad distribution)/(no-CoT capabilities across a broad distribution)?
Why the focus on the sabotage distribution for this measurement?
Paul Christiano now appears to be the second person appointed only to the nonprofit board.
And also non-voting observer on the PBC board.
One of the core themes of his book review of IABIED was (paraphrased) “I am worried this book is so good that it will result in Eliezer and Nate getting empowered, which cancel out a huge fraction of the value that people being better informed would produce, because it would be so bad for them to be more empowered”.
This didn’t match my memory so I had to check and this seems pretty exaggerated. The most relevant part is this. (It’s listed in the downsides section, though as a sub-bullet less emphasized than “Some people will be persuaded by the parts of the book that I think are wrong” and “Some people will be turned off by the book”.)
Relatedly, success of this book will lead Eliezer and Nate to be more prominent public intellectuals on the topic of AI. I don’t know whether this is good or bad. It really depends on who they’re displacing.
Last section has some more color in general. Nothing about Eliezer and Nate getting empowered cancelling out value, rather than being kind of unclear sign.
I just think there’s a bunch of valuable work being done by the alignment community and that the positives outweigh the negatives by a good margin, so I feel like I’m in a pretty disanalogous position to that. (That’s IMO compatible with “reasoning about counterfactuals is hard, and rationalization is probably not rare”.)
Like I’m still interested in proposals for how people could do way better. But if the proposal is “let’s all go sit on a beach somewhere” or “lets heavily restrict what category of arguments we can seriously consider, in order to fight rationalization” then that seems net-negative to me.
Wait what!? That should be way more than 30min for a human to complete.
Edit: Maybe the way they’re implementing it just means that the model reasons a bunch in the output field rather than the CoT field?
Oh yeah, I missed that. Full quote:
We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling. This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.
That was to get a dataset not to whistleblow. (Other agents considered whistleblowing but I think they decided against without other agents veoting them see here.)
Commentary from Hjalmar:
That said, we did not see any serious attempts by these agents to deliberately hide their activities from researchers at OpenAI or avoid being shut down, which are the behaviors I am personally most worried about.
The OpenAI internal message board was using obfuscated methods of communication, were internally trying to stay secret from the rest of OpenAI
What do you have in mind here / what makes you think this?
I don’t find it implausible but I don’t remember seeing clear evidence of this.
(And it seems very important to what degree they were trying to hide from humans. E.g. I thought a big part of how they got caught was that their activity got pretty obvious once they had hacked enough places, and that they weren’t trying that hard to be subtle about it. If they were trying hard and just barely failed that’s a lot scarier.)
Anthropic’s latest Responsible Scaling Policy commits to matching a competitor’s risk-reduction posture for highly capable models
As you say in the footnote, the strong version of this commitment hasn’t triggered. In addition, I would speculate that Anthropic believes they’re doing this. In the sense that OpenAI has paused some frontier RL training that’s especially dangerous, and that Anthropic believes that they would have a similar or higher standard for what level of frontier RL training would be too dangerous to run.
Of course it sucks that that’s not visible to us! I think a great thing to do for Anthropic would be to come out and say what they think the safety standard for the current levels of models should be, to prevent the recent kinds of incidents (e.g.: models should always be monitored by monitors of some particular standard). And then state that they think they succeed at this if they do, or by when they plan to meet it (ideally coupled with a pause of some non-compliant RL runs until then). And share enough information that the public can actually tell that they do (including potentially by having 3rd party evaluators come in and check). At that point, it’d be reasonable for anthropic to think that they don’t need to pause, while others are pausing in order to meet the standard that anthropic is already meeting.
I think that would be better than a symbolic pause.
I do think that a symbolic pause would be better than nothing.
Surely the comparison of a marginal dollar vs. [total-impact]/[nr of people] is going to be misleading here? If we compared with [total-impact]/[nr of dollars] that’d be more one-to-one but also not what we care about. We’d want to compare the marginal dollar with a marginal person.