Formerly alignment and governance researcher at DeepMind and OpenAI. Now independent.
Richard_Ngo
This is the kind of problem that motivated my thinking about aligning to virtues.
Two disagreements with you Habryka:
I interpret Zack’s “you are making false claims” as stronger than “you are incorrect” but weaker than “you are lying”/”you are deliberately making false claims”.[1] I think that a more epistemically virtuous version of Zack would have separated “you’re incorrect about two of these people” from any further inference he wanted to make about whether you did it deliberately. Instead, the phrasing he used is halfway between the two, strong enough to carry some implication that you’re lying, but not actually making that accusation outright. Given this, I think a more epistemically virtuous version of you would have tried to clarify whether he was claiming you were deliberately or just accidentally making false claims, and wouldn’t summarize the interaction as Zack accusing you of lying unless he confirmed the former.
I’d take this back if he specifically said elsewhere that you were lying about this, but I haven’t seen that. Note that in his comment above he mentions the case “if you do accidentally make false claims”, which suggests that he’s not interpreting “make false claims” as requiring deliberate intent to lie.
The claim you made was not “these people have criticized Said to me”, but rather “these people are [such people]”, which is most naturally read to be referring to Said’s description of “[usefully contributing] authors who find this person’s very presence in a discussion so ‘unpleasant’ that… it’s enough to discourage them from posting on LW altogether”. And so it might be true that people have criticized Said to you in the past, while false that they currently fall into that category overall. In other words, the most natural interpretation IMO is that their comments to you are evidence about whether they’re in this category, not determinant of whether they’re in this category, and so you could be wrong without lying.
I also think that it’s fairly obvious from the outside that when you defended your statement you were interpreting it as something like “these people have criticized Said to me”. A more epistemically virtuous version of Zack would have noted that the thing you literally said is not “these people have criticized Said” but rather “these people are [the kinds of people Said was describing]” and asked you to clarify which of these you were actually defending. Instead, Zack made claims like “(apparently putting Habryka’s own word against Alexander’s on the question of Alexander’s opinions about Achmiz)” which seems like a (motivated) failure to understand what you were trying to do.
To be clear, I find it very understandable to be frustrated by all of these interactions, since both Said and (to a much lesser extent) Zack phrased statements in ways that I consider to be norm-violatingly aggressive. Hence why I said above that I continue to think you’re doing a good job with all of this. However, on these specific points of interpretation, you seem to be incorrect or at least uncharitable, most specifically in interpreting “making false claims” as an outright accusation of lying.
- ^
I also think that “you made false claims” is slightly weaker than “you are making false claims”, because the present tense implies that there’s more of a continuous deliberate action.
My take after observing this exchange and evidence:
My sense is that Habryka was doing a kind of frustrated exaggeration when making that list, in the sense of reaching for examples that were a bit marginal (e.g. he had some evidence but not strong evidence that this person fell in this category). I notice this particular mental pattern in myself sometimes when I’m trying to prove a point.
It would have been more epistemically virtuous for him to respond to Zack’s objection by noticing confusion about how two of the authors on his list had characterized themselves differently in writing. This kind of confusion might then have led to a productive reevaluation of exactly what criticisms he was making of Said.
However, it’s hard to do the move of noticing confusion when you feel like you’re in an adversarial conversational setting, and both Zack and Said were behaving fairly adversarially. For example, “You are making false claims” is a pretty confrontational way for Zack to phrase his point, as is the “clearly contradicted” claim (it sounds like Jacob and maybe Scott were “such people” at one point in time, and then stopped being such people, which makes the situation less clear).
Re Said being adversarial—yeah, his comment above Habryka’s is extremely annoying to me, even reading it a year later. For example, describing his opponents’ position as “absurd”, expressing incredulity that people who are annoyed by him can make useful contributions, the sly dig of “none of them are remotely flattering. But such things aren’t what you have in mind… are they?”, etc.
Actually, rereading Said’s comment above Habryka’s made me significantly more object-level supportive of banning him. There’s a whole bunch of conversational moves he does in even that one comment which read to me as highly optimized to be corrosive to discourse. The thing it most reminds me of is the twitter meme “It’s amazing how much political discourse is just people pretending not to understand things, thus making discourse impossible.” Said is both repeatedly pretending to not understand things, and also peppering insults throughout his comment behind the fig leaf of asking questions. Each of those independently seems ban-worthy if he were doing them regularly.
Overall: I already had the sense that Habryka sometimes makes overly strong claims when he gets mad. This is a flaw, but given my own evaluation of Said, I think that getting mad at him is a pretty appropriate response, and so this exchange doesn’t much change my original read that Habyrka is doing his job well (at least with regard to this issue).
I think that’s just downstream of a mistake people made where they underestimated how easy it would be to build somewhat narrow systems that solve these sorts of small-horizon natural language problems
But the systems that are saturating IMO math, MMLU, etc, are not very narrow in the sense that people would have used the term a decade or two ago. So you can think about their mistake as the inability to imagine systems which have become much less narrow (to the extent that a single system can saturate almost all language and vision benchmarks from a decade ago) but are still far away from taking over the world. Then the question is: why not expect the same to happen over the next decade?
Selective Optimism: a critique of AI 2040
I have concluded over the last year that mentorship is extremely important for learning to do good conceptual reasoning.* Someone who’s sufficiently good at conceptual reasoning themselves can give rapid feedback on the specific mistakes you’re making (and often the reasons why you’re making them); also, skill at this varies by orders of magnitude.
Basically, it’s the classic finding that Nobel prizewinners did their PhDs under former prizewinners at wildly disproportionate rates, applied to a domain where success is much less legible.
* Context: my rough estimate is that I’ve both given and received more mentorship about doing good conceptual reasoning over the last year than the rest of my life combined.
Flagging that the “so I resigned myself to the world where I had to take responsibility” part makes me think you’ve updated on the object-level question of whether MIRI is succeeding, but not the meta-level question of what mechanistically went wrong at MIRI.
(I realized after writing this that maybe this type of comment won’t be very productive coming from me in particular, since we’ve had somewhat adversarial interactions lately. But I decided that it was better to just note that metacognition explicitly rather than not post a comment I otherwise would have posted.)
I meant something like “as far as I can tell, Habryka’s procedure for thinking through this decision seems to have been a reasonable one.”
This doesn’t mean I’m vouching for the procedure, because there are many parts of it that I can’t verify/don’t know about.
But it also doesn’t just mean I’m deferring to Habyrka, because I am trying to figure out if he’s running a good decision procedure, and would have commented differently if I thought he wasn’t.
It’s something more like “Habryka has a job, and he seems to be doing his job in a reasonable way in this case”.
Zooming out: on the meta level, the reason I revisited this comment is that I saw some of your exchanges with Cade Metz, in particular the ones where you mentioned that you were engaging with him because you believe in Speech so strongly that you’re willing to pursue a quokka strategy. To me, this seems both admirable in many ways, and also like a fairly unreasonable approach (as I believe Michael also mentioned in the post). Given this, by default I don’t plan to engage with your “Comment on “Banning Said Achmiz”″, because I’m worried that you believing in Speech to this extent will end up being an unproductive crux for us. However, I’m open to arguments for why I should change my mind.
I think it was somewhat true in 2016 but also most AI safety people weren’t very aware of it; IIRC the ostensibly-safety-motivated capabilities researchers at OpenAI in 2016 weren’t widely known to be such.
Preregistering a disagreement with my MATS fellow Erez Abrams:
He thinks the right number of levels of agency is 2. I think that 2 is too large, and the right number of levels of agency is 1.
(This isn’t mean to make much sense to anyone else, but hopefully one of us will end up obviously right at some point, and will owe the other person a poem about the beauty of their conception of agency.)
Similarly, if you told an EA from 10 years ago that a non-negligible number of EAs were doing capabilities research at the leading AGI lab, they’d also assume something had gone very wrong.
Hmm, in some sense I’m on board, but also I have an intuition that one of the “problems you’re trying to solve” is something like “getting trapped by your own beliefs”, so a high-dimensional belief web also creates more scope for that.
I expect that two or three actual examples would be super helpful in conveying the texture of this problem and how much I should care about it.
Great comment, ty.
Traders in Garrabrant markets are perfectly capable of doing reasoning on logical statements about Garrabrant inductors
Yepp, agreed. Garrabrant induction feels big/rich enough that it wouldn’t surprise me if we “found” a bunch of complex structure in how traders interact mediated by their reasoning about propositions about each other.
However, it does seem like the details of how traders interact might be important. For example, a trader can have a policy of doing the opposite of what the market says it’ll do if the market becomes too confident, which limits the resolution at which market propositions can “see” the traders.
Whereas Scott Garrabrant has suggested to me the possibility that you could just let traders “see” each other’s trades directly before they happen, then formulate their own trades as a function of those other trades (though figuring out what trades actually happen would then require finding the fixed point of all those functions). This seems like it’d be meaningfully different from the existing setup.
If each agent delegates a chunk of goal to a sub-agent, then the chunk of goal which can be delegated is limited in size by the capabilities of the sub-agent: too much goal and the sun-agent breaks down.
Ah, interesting point. This also basically seems right. I’m reminded of a comment I saw on twitter that one way you can harm someone is by giving them goals that far exceed the scope of their agency/control.
Agents as Webs of Beliefs
A different Deutsch theory, but it’s not really a theory, more like his gesture towards what a theory of everything should look like. More here.
Epistemic status: wild speculation
There’s an interesting line in the old MIRI paper on Definability of Truth: “Tarski’s result on the undefinability of truth is in some sense an artifact of the infinite precision demanded by reasoning about complete certainty.” This feels vaguely analogous here to the uncertainty principle in quantum mechanics, which also prevents infinite-precision knowledge. (Has anyone discussed this before?)
The connection jumped to my attention because I’ve also been thinking lately about the concept of causation in physical and logical universes. Intuitively speaking, the concept of Pearl’s do-operator feels pretty sensible in our physical universe because there’s a bounded lightcone which an event can affect—so that when we hypothesize taking an action, we can mostly ignore the implications for things in the past, or for simultaneous actions by other agents. (See this old post by Yudkowsky for more.)
FDT generalizes the do-operator to the realm of logic, which doesn’t have either timelike or spacelike separation by default. There’s no privileged direction such that some logical facts affect others but aren’t affected by them. And there’s no notion of “distance” between different logical facts—e.g. any two facts A and B can be combined to give a proof of A AND B.
It’s possible that we could ground some version of logical timelike and spacelike separation using concepts like “what order does some proof system generate these facts in?”, “how hard is it to prove one from the other?” or “how much computation do you save by calculating them together vs separately?” But I don’t currently see much promise in this, since it all seems so representation-dependent.
An alternative idea (inspired by some of davidad’s thinking) is that intelligent minds might create logical timelike and spacelike separation as a consequence of their decisions. Demski’s conception of logical time (which was anticipated by Lacan!) is mostly motivated by the idea of agents recursively reasoning about each other. It seems like such agents could create a kind of directionality by simply refusing to reason in certain ways, as part of a correlated game-theoretic strategy. (As an aside, the idea of refusing to do certain kinds of reasoning about other agents seems like an important building block for defining concepts like respect, trust, and faith.) This could lead to a logical analogue of the lightspeed limit which is less of a hard constraint, and more like a highway speed limit: breakable, but in a way which might incur larger consequences (specifically, exclusion from the coalition that coordinates itself by imposing spacelike and timelike structure in the logic of rational minds).
More generally, I want to coin a term for the idea that there are strong parallels between the natural sciences and the mathematical/cognitive sciences at many different levels of abstraction, such that they could eventually be unified into a single theory of physico-logical phenomena. This is related to Deutsch’s concept of a non-reductionist theory of everything, but he doesn’t (to my knowledge) point specifically at the idea of parallelism between these two domains of inquiry. Meanwhile Friston’s work on the free energy principle attempts to draw a link between cognition and thermodynamics, but (again, to my knowledge) he doesn’t hypothesize that such links arise at many other levels (quantum physics, chemistry, biology, etc). For now I’ll call this the physico-logical unification hypothesis, but I’m open to better suggestions.
Ooops, good catch. I was going to say “because of this, you can’t assign credences to them which sum to one”. But I’m not sure that this is quite correct—I think there’s a deeper reason why it’s hard to assign credences to them at all (because traders/theories are only ever approximately correct). So I’ve just deleted the sentence fragment.
Early last year Longview Philanthropy invited me to attend an AI safety fundraising dinner they were hosting for Jane Street traders in Hong Kong as an expert guest. In the lead-up, I got on a call with them, and gave my honest opinions about their funding recommendations.
I summarized these opinions in a follow-up email to them at the time as “Overall I’m very excited about three of them ([redacted], [redacted], [redacted]), and moderately excited about most of the rest. But I’m pretty skeptical of some of the policy advocacy (particularly [redacted] and [redacted]).” (Though it’s possible I came across as more pessimistic in the call than I sound in this summary.)
As a result, they rescinded the invitation.
It’s not clear to me that they were wrong to do so, because my sense is that it’s normal to prioritize bringing more sympathetic experts to such events.EDIT: Upon reflection, even if this would be normal in most philanthropic contexts, I want to hold organizations associated with AI safety to a much higher standard. I expect that selection effects like this one add up over time to give funders a distorted view of the ecosystem, and IMO a more virtuous fundraising process would find ways to mitigate them—e.g. as the AI Futures Project did by asking me to release a critique of their scenario. I expect this kind of virtue would, if widely adopted, add up to make the AI safety ecosystem much healthier. (To Longview’s credit, they were interested in talking further about my criticisms privately, which I didn’t end up getting around to due to busyness.)I wanted to share this anecdote primarily to give people a better sense of the kinds of dynamics that block information flows in the AI safety community. As another example of such dynamics, I’ve redacted the names of the charities from my quote above because Longview asked me to treat their recommendations as confidential at the time (and as far as I can tell, they’re still not public). However, I’m pretty skeptical that this kind of secrecy is healthy for a philanthropic organization.
EDIT: since I’ve edited the text above to take a stronger moral stance, I also want to clarify that I didn’t hold my current opinions about the importance of transparency at the time, and called Longview’s decision “very understandable” in an email to them. So I don’t want this post to convey “I was wronged by Longview”, but rather “we both participated in a dynamic that with the benefit of hindsight I consider to be unhealthy and worth describing publicly”. I should have expressed that at the time, and if I had they may well have changed their minds.