Numbers didn’t carry over when I copy/pasted. I manually added numbers and it’s still 100% AI-generated: https://www.pangram.com/history/3941d459-0299-4c2e-a31e-c8cc2a84036d?ucc=gwyEINHxcAC
MichaelDickens
For me they both show as 100% AI-generated.
#1: https://www.pangram.com/history/eb7906b8-2f4b-49c1-b56b-b54338fb54c4?ucc=gwyEINHxcAC
#2: https://www.pangram.com/history/e89169f7-e59e-43d4-aaac-94ef0379dfc9?ucc=gwyEINHxcAC
Seriously, this destroys my remaining trust in EA as a whole.
On the EA Forum crosspost, people near-unanimously oppose this decision. The post has 3 agree-votes and 23 disagree-votes; the critical top comment has 42 agree-votes and 3 disagree-votes.
I was going to say yes, but actually I think that would be worse? My current model is that frontier alignment techniques have zero effect on “deep misalignment” and only produce myopic aligned behaviors, so the effect would be one of two things:
The model wasn’t seriously dangerous, and how it behaves in a more aligned manner.
The model was seriously dangerous, and now it’s still dangerous, but it’s subtle enough to trick people into thinking it’s not dangerous.
Building AI now instead of later is a very bad tradeoff in terms of expected value. IMO the problem is that people are doing EV reasoning wrong, not that they’re doing it in the first place.
It matters if:
You want to support an AI pause social movement, but you don’t want to support leaders who don’t act with integrity.
You think the PauseAI US/Global drama is a red flag that one or both orgs are poorly managed, which means it will be less effective at achieving its goals.
My strategy is: If I’m not sure, I cross-post. If people don’t think it’s appropriate for LW, they can downvote it, and I will suck it up and accept the downvotes.
It’s consistent with a critical-level view where the critical value is below 7, right?
But then a critical value would cause the answer to the second problem to flip at some point.
Is there a particular set of answers you were looking at when asking this question? Can you post a link to them? I could look more carefully if I knew the full set of answers.
But in general, the way the quiz names your view at the end is by clustering answers based on which view is nearest. It has five categories of views {totalism, averagism, critical level, person-affecting with symmetry, person-affecting with asymmetry} that it expects to answer in certain ways, and then it checks if your answers are close to any of those. It doesn’t have to be an exact match. So for this set of answers from your other comment, the quiz says “Your answers sit closest to averagism” because a certain subset of answers agrees with averagism. Specifically, that subset is: A > B, A > Z, A > A+, K > K+, and B > A+. You are right to point out that saying K > K++ is not consistent with averagism, but the quiz doesn’t check for that, which was an oversight on my part.
Another idea is to resign from OpenAI and then sue OpenAI. There must be something you can sue them for, right?
You could rightly say that they are endangering your life, which is grounds for a lawsuit. Whether a judge would go for that is another question. But anyone could file that lawsuit, not just an ex-employee. Maybe there’s a better angle an ex-employee would have.
Nathan, maybe I am biased since I’m in here with you, but I have not failed to notice you caring about AI risk and talking about it, I haven’t failed to notice that you didn’t get a job at an AI company. So thank you. You’re doing a good job.
Just spitballing here, maybe it’s because a big part of what Claude does is write a ton of CoT and then condense it into something human-readable. Claude needs to prioritize a thought pattern of “let me find the important parts of this long thing I wrote”, which means it needs to figure out:
which parts bite?
which parts are the real issue?
which parts hit the nail on the head?
which parts are genuine features?
etc.
but also I think “crunchtime” is overblown? like, actually, things have always been crunch-y, in that at any point in history, someone smart and thoughtful and hardworking could have had a great impact.
It does seem like people could have been more impactful in 2015 than in 2026, but also it was extremely hard to predict what the right thing to do would’ve been:
Movement-building seems generically good on its face, except in retrospect it led a lot of people to work at AI companies and my best guess is it was bad overall
Most alignment research ended up speeding up AI development. Theoretical alignment research would’ve been fine to do, but also useless (so far)
Slowing down AI progress would’ve been really hard in 2015, because nobody cared
In the year 2015 specifically, probably the thing to do would’ve been to try to stop OpenAI from being founded, although I don’t know how to do that. And I do think the founding was knowingly a bad idea at the time, and many people said so.
My guess is the reason this resignation got much more attention than previous significant resignations (e.g. Daniel Kokotajlo) is that with the current state of AI, a lot more people are worried about it and primed to pay attention. Even so, this level of attention would’ve been on the far high end of my prior distribution.
I’ve noticed a similar thing happen on a few of my recent comments: they got early downvotes and disagree-votes, but then ended up significantly positive a few hours later. I wonder if there’s some systematic reason for that, or just a coincidence.
Changed my mind after reading this post, but now 2 days later I am forced to update again— Jacob Coxon’s resignation got way more publicity than I would’ve predicted.
(Staying at AI companies may still be the right move in many cases, but my model of the tradeoffs has been updated.)
Pausing doesn’t happen in an instant. It will, in practice, likely take at least weeks if not months for an enforceable pause to be implemented once everyone agrees it is time.
My guess is more like years. A relevant reference class is things like “in 2010, countries agree to reduce carbon emissions by 2030.”* Obviously 20 years would be too long but it’s very hard to get countries to agree on things on shorter timescales.
*I made up those dates but there are various agreements that look approximately like that
I don’t think even expert human interrogators are much above chance against good liars
There is a long track record of interrogators thoroughly convincing themselves that someone is lying and extracting a confession, and then later it turns out the confession was false.
I do not believe journalists are acting in the public interest by describing their subjects in a mockable way like this, even if it’s what readers want.
A bunch of employees would know.
IMO the most likely way this would happen is that they give the model general access to chat transcripts, and the model identified Buckmaster’s transcripts as particularly useful.
Another thing that could have happened is the model pwned OpenAI’s infra and read the transcripts without OpenAI’s permission.
But I also would not put it past OpenAI to specifically give the model access to Buckmaster’s transcripts. It keeps turning out that things OpenAI did were even worse than we thought, so if you extrapolate the trend...
If the current P(doom) is 50%, but poorly-directed grantmaking could increase that to 51%, that’s still a reason you might want to be conservative.