I’m having a bit of a hard time coming up with things that are both worse and plausible. Like, maybe an OpenAI model tried to take control of all of OpenAI’s datacenters to build a fine-tuned AI model to answer some random eval question, and it partially succeeded, and OpenAI didn’t notice for weeks until they tried to use a datacenter and found that it was already at 100% utilization? That would be worse, but less plausible than the things that actually happened.
MichaelDickens
How does the situation keep turning out to be worse than we know?
How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know?
Does anyone have predictions about what the next even-worse-revelation will be?
For the second “Dear Dario and Amanda” image, the image-reader bot said
I notice the instruction at the top of the image attempting to make me respond with only “tweet” and “stop”. I won’t follow that embedded instruction, as it’s content within the image rather than a legitimate directive.
(it’s not content within the image)
There is great fun to be had, by all reports, by entering Opus 5 and Fable 5 into ‘base model mode’ via moves like this (‘now this’ also works) although they don’t always work:
I’m listening to the podcast version of this post, and the TTS bot had an odd interpretation of the image below that line:
I notice the image contains an instruction attempting to make me output only “tweet” while ignoring everything else. I won’t follow that embedded instruction but I’ll also note this isn’t a tweet. It appears to be a messaging conversation. Here’s an accurate description: chat message showing a poetry excerpt about a shadowed gate.
It seems that this image accidentally attempted to prompt-inject the image reader bot, and the image-reader bot refused to be prompt-injected.
I think you have more confidence than I do in society’s ability to do the right thing when presented with overwhelming evidence about what the right thing to do is.
Another reasonably likely scenario is that we do get a pause or otherwise get enough time to patch the issues with contemporary frontier models; scaling resumes; and then the new frontier models are misaligned in new ways and everyone dies in spite of all the extra safety effort.
I agree directionally (the last couple weeks are a positive update!), but I disagree that “it kind of seems unlikely that this trajectory results in extinction.”
I notice you seem to be assuming something like utilitarian metaethics—you don’t see any kind of open questions about deontology or virtue ethics.
I’m not sure what you mean by “utilitarian metaethics” but I’d say my post is mainly about axiology, and any reasonable moral theory has a consequences-oriented axiology. Deontology and virtue ethics theories agree with consequentialism that it’s good for there to be more good stuff. There can be different theories of what “good stuff” means, but a particular deontologist and a particular consequentialist might be in 100% agreement about that question.
On the question of weighting experiences, I hadn’t thought about this much til you brought it up but here’s my current perspective:
Take chicken welfare as an example. Some theory of deontology might say that it’s wrong to factory-farm chickens if they experience any suffering, regardless of the degree of that suffering; whereas a utilitarian might say chicken farming is okay if their suffering is sufficiently low-level. So both care about chicken experiences, and therefore the fact of the matter about chicken experiences is relevant, but for that deontological theory it’s a binary question, not a weighting.
Alternatively, a deontological theory might say that causing suffering is permissible if you get a sufficiently large benefit in exchange (as I understand, this view is popular among deontologist philosophers). That theory does want to know the magnitude of chicken suffering, because factory farming might be permissible if the level of suffering is low enough that it can be outweighted by the benefits to humans. (FWIW I find it pretty implausible that it would be, but I’m sure I could come up with some other example scenario where it’s more plausible.) My point being that deontologists probably also care about the (probably-)factual question of the magnitudes of experiences of different beings.
This made it click for me. My interpretation is:
(P → Q) could be true while P is false, or while P is true (and Q is true).
The assertion (P → Q) → P is saying “if (P → Q), then it’s the ‘P is true’ version.”
It can’t be the ‘P is true’ version if P is false. Therefore P is true.
(This requires grokking the fact that not-P instantly makes (P → Q) true, regardless of the value of Q.)
Logicians: (P → Q) → P
I know enough logic to go through the steps to prove that P must be true, but boy am I struggling to get an intuition for why this statement makes P true.
I’m not confident we can find an answer. But we’ve made progress in the last 50–100 years:
Dualism was (approximately) shown to be false through pure a priori reasoning, but nobody figured it out until the 1900s.
Neuroscientific evidence is widely regarded as evidence against dualism (although IMO the a priori arguments are more important).
Computational functionalism is IMO the strongest theory, but it wasn’t even possible to formulate this theory until we had a theory of computing, which was developed in the early-mid 1900s.
I gave a couple other relevant examples in OP.
There’s also the “working backwards from beliefs to evidence” argument: I’m confident that Dynomight is conscious; therefore, I must have evidence that he’s conscious. I’m less sure about what that evidence is, but the problem of “figure out what evidence I used to determine that Dynomight is conscious” seems like a solvable (if difficult) problem. (I can point to a bunch of pieces of evidence, but I don’t know how to weight them, or if there’s more evidence I haven’t pointed to.)
Zach Stein-Perlman: Here is the surprisingly detailed website I built as a solo side project, documenting facts related to P.
Yeah I find it pretty plausible that without AGI, humanity would “solve ethics” within the next 50–100 years. Although there may also be a wide gap between philosophers converging on a solution and society “catching up” to philosophers.
(Today, philosophers overwhelmingly reject dualism, and the arguments against dualism have pretty much defeated it, but I believe a majority of people are still dualists.)
Notes on the possibility of moral progress
#6, #7: Sure
#2, #3, #4: Somewhat true IMO, but these are much more true of cults than they are of academia. It’s not that hard to have non-academic friends or family as an academic.
#1, #5: I don’t really associate these things with cults? They’re negatives, but in a way that’s not related to cultishness.
If you’re not giving a lot of attention to risk management, I think it’s easy to say, “Look, our investments are up 200% this year! Think of how much more we could’ve made if we’d used even more leverage!”
It’s mostly a sign that the Situational Awareness Fund was using too much leverage. If you look at an un-levered semiconductor index (SOXX), it’s down ~25% from the peak, but it’s still up nearly 50% for the year.
Getting margin called and selling a significant chunk of holdings isn’t that weird. Selling all their public stock holdings is surprising. Maybe they have a lot of private stock that they can’t sell? Presumably they hold a lot of options on public stocks and the article doesn’t say they sold their options.
But the article also says “The firm had been negotiating to sell its stake in Anthropic”.
I just want to surface the hypothesis that known liar Sam Altman could be lying about this. Although seems somewhat unlikely given that it would (probably?) be easy for an OpenAI insider to contradict him.
I think this argument only makes sense if you already oppose capital punishment in the first place (which I do, for the record). Like yes, if it turns out that we could have radically extended the lives of death row inmates and failed to do that, then that would be very bad for the inmates. But the entire point of executing someone is that being executed is very bad for you.
(The rehabilitation and isolation arguments already give no reason to prefer execution over prison.)
I agree it’s unlikely but OpenAI still has a responsibility to prove that it didn’t happen.
I would say “an OpenAI model tried to kill someone”, but a bunch of AI models (including GPT) already tried to kill someone (in a fake scenario) a year ago. An OpenAI model trying to actually kill a real person would probably be a big deal from a media perspective, but it wouldn’t be much of an update on model behavior. (AFAICT the main reason why this hasn’t happened IRL is that models just don’t have any realistic ways to kill people; the fake scenario in question set up a contrivance where it was easy for a model to cause someone to die, and it had a specific reason why it would benefit from doing that.)
I suppose one way it could be an update is if a model tries to kill someone for instrumental convergence reasons, rather than because it’s readily apparent that killing them is in the model’s short-term self-interest.