Like, you’re just watching the externally visible actions of the agent? I think this would in fact not really pin down the answer, except to, like, Solomonoff induction (basically reading / guessing the internals of the agent by induction, and then evaluating those internals). There’s definitely no such thing as a “dangerous action” in this context, in the relevant sense. I mean, your question kinda makes sense if you ask it like “Is it more important to go from no judge of alignment to an expensive judge, or to go from an expensive judge to a feasible judge?”. But the thing about actions makes the question not make sense as stated.
TsviBT
It’s really not (but if you won’t think about it more then I guess there’s not much more to say). At least naively, if you rephrase the question to include the full state, then the question is basically “can you tell whether the mind is aligned or not”. Did you mean to ask something equivalent to that?
It seems like mostly a type error to ask if an action is dangerous, or the wrong question. Actions are part of a whole life story of behavior of a mind. There are actions that are definitely dangerous on their own (“press the nuke button”), but all the useful good actions are ambiguous with dangerous actions, if you just look at the action on its own. E.g. “get much smarter” is a dangerous action if the surrounding mind is dangerous, and is good if the surrounding mind is good.
Mainly I meant to say “you probably should have substantive uncertainty about the extent to which it’s a sign about their capability, because (I assume) you don’t know the sorts of people they’re trying to hire and how those people orient to politics and whether that’s a good orientation or not”.
Talking to Richard and hiring Richard are two different things. It’s plausible to me that all these things are true simultaneously:
Richard would do a great job at contributing useful research directions / thoughts / etc.
Richard would be great for the org culture and treat everyone well / respectfully / etc.
Richard’s politics have substantial admixtures of attitudes that are dumb (both incorrect and harmful), and indicative of general attitudes / thought patterns that are reasonably taken as hostile / toxic / etc.
Other good people have low thresholds for that sort of indication, for good reasons, and in fact would be less likely to join and/or fund the org if Richard is there
The expected net effect of hiring Richard is negative, including integrity / upholding principles (it’s not necessarily unprincipled to be like “I think Richard’s ideas are good enough that it’s worth some tail risk, but I get that lots of other good people wouldn’t have that information, and furthermore have good reason to distrust him for political reasons, so I won’t hire him”)
I would not guess that all of these are true, but I wouldn’t be so confident without, for example, having talked to several potential hires for the org.
Well, in theory it could for example suggest, directionally, whether it’s more a “my rationalist tools are breaking in my hands” situation vs. a “I have failed to live up to my rationalist tools” sort of situation.
I think it’s quite difficult to kill all humans / forcibly permanently disempower all humans. So yeah, to answer your original question, that’s probably a crux for something.
If humans decide to disempower themselves, sure. That seems a lot slower and more reversible than a scenario where the AI systems are actually preventing humans taking back power.
Re/ MtG, that specific question has occurred to me as interesting re/ AI capabilities. But I don’t really know enough to figure out how hard it is. I would not especially bet against “an AI will win a new draft format tournament by 2035” or something.
(That document seems to have been largely AI-written, based on sampling paragraphs from ~3 spots to Pangram.)
A general issue with the intense versions of Buddhist practices is that very often they are aimed at detaching / deconstructing / destroying / disavowing various elements of one’s mind / agency / values. Often these are important elements, and don’t necessarily automatically stay intact / regenerate themselves. This can have various good consequences, but also obviously by default has a bunch of bad consequences. Maybe there are good ways to do that (for example, by giving these elements space and other resources needed to reconfigure / regrow / reinterface / etc., in an appropriate dovetail with deconstruction). But my overwhelming impression, admittedly in most cases from shallow exposure, is that these practitioners either do not understand what they are doing, or think that this is a good thing to do, which it very obviously isn’t (similar to how suicide is very obviously not a good thing to do, with similar exceptions). (In defense of my conclusions even given shallow exposure: when I have tried engaging a bit more deeply, the background assumptions seem to be blatant, explicit, and strongly conditioning everything else; and the fruit of the tree seems often bad, ~never super impressive, and only mildly good when done in a non-standard, non-pressured, a la carte way.) Spiritual and social pressure towards these practices seems to be very often carried by unspoken assumptions of the obvious desirability of these things, preying on people’s desire for relief / escape / growth / evolution / self-understanding / calmness / etc.
(Not accusing you in particular of any of that, and I haven’t engaged with your work; it’s plausible to me you’re doing a thing that works well for you, and that others have similar results. But I do want to point at intense versions, including probably “Westernized versions of Buddhism that have had the supernatural stuff stripped out of them”.)
As a prior, people often offer a false dilemma of “you gotta sell your soul because our cause is Good and Just enough.”
Indeed, this is one of the central engines of abuse / cults. It’s spiritual pyrite: You identify what the mark deeply wants; present yourself / your group as providing it; and then when the mark has any thought patterns you don’t like, you frame those as being against their goal of getting what they deeply want.
The use-cases of robotics beyond current robots, especially home robotics require handling lots of environments and have varied challenges, so it stress-tests AI generalization.
“it stress-tests AI generalization” But does it though? What does it stress test? That seems to assert a generalization from AI generalization about XYZ robotics tasks to ABC intelligence explosion tasks (or something along these lines?). What’s that generalization / what justifies it?
Unlike other domains, there’s no unexploited algebraicness/overhang that would allow AIs to be useful for these cases without generalizing, because if there were such a thing, very limited current industrial robots would already have been used. They haven’t, so there’s no room for algebraicness/non-generalization to be an issue.
I don’t understand this. Industrial robots of course are used very widely? Or I guess you typoed, and you’re trying to say, LLMs or other AIs would have been used to compute very non-obvious complex actions that take advantage of algebraicness in the narrow domain? Or you’re saying something about humanoid robots, or about softer / less rote tasks (laundry folding rather than installing a chair in a car)?
For these reasons, if current AIs improve on robotic tasks, it is a much stronger signal that they’re generalizing, and therefore is a signal that AGI could come in 10-20 years, and depending on how fast they improve, this might shift to 2-4 years.
How do you get from the qualitative and relative statement about robotics being “a much stronger signal” to numbers and years?
Or to put it another way, you seem to treat AGI as though there’s a discrete set of challenges to get there, one after the other to be solved, while I instead treat AGI as a set of challenges which all have continuous metrics, and in particular there’s ways to reuse previous progress, because the load-bearing elements of AGI depend less on a particular paradigm/theory of intelligence and they depend more so on compute.
My guess is that these things are not actually cruxes? Not sure. Instead I’d guess that the cruxes are more simply
the quantitative amount of insight (regardless of continuousness) remaining to get to AGI
the degree to which those insights are blocked on difficult thinking that current AI doesn’t accelerate by much
Thanks for A2A. I think it’s going to be hard to communicate about this because there are fundamental conceptual boundaries. It could help for you to write out the basic structure of the argument in simple concise syllogism format; and then expand on background enthymemes involved in the concept used for induction. This is to help address objections I might make about “wait does this concept actually support this inference / induction”? (Cf. https://www.lesswrong.com/posts/i7JSL5awGFcSRhyGF/shortform-2?commentId=fXtNjxLevGaiJHSyj )
In particular, in this case you’re using some concepts here:
they have a positive but subhuman amount the property of effectively and continuously acquire deep knowledge and exploit this knowledge to construct and execute goal-directed plans over long lifetimes and consequence horizons/generalization.
And you’re doing an induction about it. And I’m just going to object that the way you’re using the concept doesn’t support the inference you’re trying to make; or at least, AFAICT you haven’t explained the concept and how you’re making the inference well enough to justify the inference.
The particular observational evidence you adduce seems interesting, but also I don’t have enough context to know what it means. (And I’m not really trying to evaluate it, because busy and because the disagreement seems to stem from the above, not from some specific observation. There are observations that would invalidate substantive chunks of my model (cf. https://www.lesswrong.com/posts/sTDfraZab47KiRMmT/views-on-when-agi-comes-and-on-strategy-to-reduce#comments), but generally when people bring in some observation the disagreement is how they are interpreting it, not the observation itself.)
Yeah, I think it should be much less hard for people not already embedded in working on AI; just seems kinda related.
E.g. Musk seems like mega-biased toward taking action, with bad results (responsible for like a quarter of frontier AI lab stuff, in some sense). The politician’s syllogism bites real hard when accelerating bad stuff is 1. convergent and 2. much easier than accelerating good stuff.
(This is probably not helpful, but just want to note that I’m interested, in a “hobbyist” sense, in the problem of the general kind of confrontation you’re mentioning, under the moniker “confrontation-worthy empathy”; if you happen to want to chat about it I’m interested, e.g. to consider different ideas for making the process more wholesome and similar.)
(I’ve been “out of the game” for a couple years.) I would append to
you might think AI alignment is a computer programming problem but it’s more of a math problem
a follow-on:
you might think AI alignment is a math problem but it’s more of a philosophical / conceptual / mental-phenomenology problem
(I hesitate to use the word “phenomenology” because it will definitely be misunderstood by ~everyone, but there just isn’t a better word; an internal phrase that was floated was “core theory of mind”, and I mean to gesture at “mathematico-introspective investigation of core theory of mind”.) This is elaborated on (cryptically / elliptically, but you could read slowly & think) here: https://www.lesswrong.com/posts/TNQKFoWhAkLCB4Kt7/a-hermeneutic-net-for-agency
In this context, the long and short of it is, I don’t think you can feasibly point people at the parts of the problem that matter.
I’m just such a special genius that only I could possibly understand the real alignment problemFor some reason people don’t seem to be inclined to plant difficult questions and let them grow over time, sit with unresolved questions, look at big things and keep their bigness firmly in mind while also somehow making cumulative progress, overhaul key elements or at least keep them firmly provisional, keep staring outside of the streelight’s glow for years, and so forth. (Something something Grothendieck something something Peter Scholze something something staring at conceptual foundations.) Or maybe it’s the thing about security mindset (Yudkowsky), or overconfidence (Dai), etc. Or maybe we’re just not high g enough. Or all of the above or something else, IDK. But again, it just doesn’t seem to work to point people at the problem. Or maybe I’m deluded, who could know.If you take all the alignment research, insofar as I’m aware of it, in the past 2 decades, and multiply that by 3, it still doesn’t even come close to solving alignment. This is very rough and grim, but it’s what I think.
I’m not sure I buy that most mathematicians will be out of a job. Maybe some of the number theorists, analysts, and combinatorialists? Or something, IDK, I’m making that up. (Based on what fields vaguely seem to have some niches for people doing mostly high-algebraicness work.) Or maybe they will be.
What should they do instead? IDK. Human intelligence amplification is still insanely underresearched, and can definitely use more brainpower. I’d be very happy to be connected with smart motivated scientists / mathematicians who may want to work on empowering humans!! My gmail for that would be: tsvibtcontact
I’m not saying it’s slop or bad advice, I’m saying I wouldn’t be able to tell, and I expect other people also wouldn’t be able to tell without effort.
I’m not aware of any of these having happened, are you? (NB: I neither up nor downvoted.)