I wonder if any strategies of this sort are pursued at CAISI (where Paul Christiano is Head of Safety).
neo
Is anyone doing anything (on a personal level) about the possible impending hacktastrophe / collapse in cybersecurity resulting from capable AI models and swarms?
I’m fairly aware of most standard advice, but curious if anyone has taken more drastic measures.
Could it be that polyamory selects for people genetically predisposed to feel jealousy less intensely? Maybe due to some degree of inhibited oxytocin signaling?
Human beings continue to be the “final node” even when AI is used for research/preparation, which I think is particularly relevant in an activity where speaking ability and charisma play a large role in competitive outcomes.
In other words, I think debaters will always be bottlenecked by their ability to articulate the ideas they have, whether or not they are AI-generated.
It seems like this will have a negative effect on debaters whose strengths are research or quick-thinking; good speakers who can effectively and consistently mask AI-usage may be able to compensate for what they lack in other skill categories.
But I think the solution is simple—competitive leagues of any sort who care about avoiding total disempowerment in their activity should outright ban AI usage and rigorously enforce such a ban.
On the other hand, I’m excited about AI-assistance in the judging process; I think determining a fair and systematic judging process/criteria is a fascinating epistemic question that can expose where human biases play a role, and inspire real progress.
Cool to see we had a similar idea, and thanks for linking my post!
The custom query feature is super useful and a clear improvement over my version (I only have it on my personal setup to avoid racking up API costs for now).
I’ve been considering expanding the project to a full plug-and-play system where users can hot-swap embedding engines, sources included in the corpus,[1] etc. - just haven’t gotten around to it yet.
I genuinely think semantic search is useful/impactful enough[2] to justify becoming a core LW feature and I really do hope to see that happen soon.
- ^
Wikipedia pages, abstracts of research papers (something like connectedpapers), etc.
- ^
LessWrong in particular has an abundance of valuable (often cutting-edge) information/thinking on a diverse range of topics, but “situating” potentially novel ideas with current (quite imprecise) search methods is very difficult, and IMO a potential bottleneck to progress.
- ^
I’ve always had an inkling that frontier labs would make increasingly rigorous safety communications as the threat became more salient. After all, most lab leaders seem to have recognized x-risk way before the AI boom and are presumably interested in continuing to exist.
And in general, I think a lot of the safety community tends to extrapolate from current lab safety attitudes and not price-in changes to communications and concerns that seem likely as capabilities, and accompanying unease, grow.
But I agree, this doesn’t feel satisfying for some reason. Was it necessary to downplay/gloss over safety and scare the shit out of everyone concerned about this sort of thing? And if it was, for political or financial reasons, it means there exist incentive structures working against transparency of knowledge/belief, which seems bad overall and for the long-run.
Side note—your link requires a login, not sure if that’s intended.
I think he somewhat answers your point here:
And as a result, our sort of meta strategy involves essentially roughly forecasting, maybe not formally forecasting, but roughly in our heads having some sense of what things are going to be potential problems over the next some time period — call it three months, maybe longer, maybe shorter, who knows — and identifying what sorts of considerations might become quite important during that time given the capabilities that we expect to have, and making sure that we’re prepared for those. And if we’re not prepared for that, then possibly slowing down, or pausing development, or talking to governments, trying to do advocacy.
But importantly, we’re not really trying to forecast arbitrarily far into the future all the problems that are going to arise with AI development. The goal isn’t “Know how to align ASI, or else do nothing.” Usually I think of us as looking at a time horizon of, it depends on which particular thing we’re doing, but often somewhere between three months and five years.
I interpret this as: we are focusing on short-term planning and ready to support a pause if we see imminent danger.
In other words, he probably is concerned with the alignment worries you mention may emerge in the future, but doesn’t consider them proximate enough to warrant present-day commitments.
I think it would be good to have easy to click examples on the welcome page, so we can get a feel for it without having to pick a post to test.
Done! Chose some high-karma posts that seemed particularly interesting for now.
I do intend to open source the code, and I’ll definitely try out the connected force graph.
I built a semantic search engine for LessWrong
As someone whose worldview was upended over the last few years because of AI progress, this post resonates. Sometimes the situation we are in just seems absurd – like, all I can do is laugh, shrug, and then go back to what I was doing. And I think this is an emotionally healthy response. Sometimes the best reaction to reality is an unbothered acknowledgement of it’s absurdity.
But I do worry that speaking of doom like this – as if it is nearly-certain[1] – is counterproductive.
I think everyone should invest time in mentally preparing for an uncertain future. Finding a way to be okay with the possibility of disaster, while staying motivated to work hard to avoid it.[2]
This is important whether or not you think doom is likely. In all worlds, those who can be effective regardless of their perceived odds of success, are the ones most capable of succeeding.
I have personally managed to find a healthier relationship with the world; retaining some whimsy and optimism despite a sober reckoning with reality. And I feel like I can actually achieve things now, as a result.
I want others to experience this as well. The ideas have already been discussed, but emotionally internalizing them takes a while, and it certainly did for me.
I hope we find better ways of communicating this so that all of us could get better at achieving our goals 🙂
- ^
When the emotional advice you give is downstream of that prediction. Like, Dying with Dignity or “Life may be very short. So make the next few years the best ones.”
- ^
I’ve been thinking about this comment a lot, although I can’t attest to any of the specific recommendations.
- ^
Apologies for the delayed response.
I should note that the post was somewhat hastily written – I agree that my categorization was not comprehensive, and yours is probably better.
I was mainly trying to point at a dynamic I see often online where influential voices present arguments regarding AI risk that completely ignore years of back-and-forth discussion on similar topics, but whose positions are interpreted as the “forefront” of the debate – leading to offshoot discussions that again, miss years of relevant literature and discourse. I think this leads to, among other things, Eliezer frequently “losing it” on X/Twitter over people apparently misunderstanding something he wrote about in detail 20 years ago.
Community Notes on X/Twitter are sometimes regarded as a major improvement to collective epistemics, but I think there is a lot of room for improvement with tools that “situate” current discussions within previous ones.
But yes, I am describing a somewhat vague and imprecise problem here, so it may be difficult to categorize or pin it down with certainty.
I think the disagreement stems from a lack of specificity on my part; ignore the specific description of the categories.
Probably, you are in soldier mindset yourself about this very issue.
I hold beliefs on it, sure. I am now interested in seeing if they reflect reality, and learning why/why not. Is this mindset inadequate, and what would make it more rational?
Separately—do you think there is promise in tools of the type I describe to combat soldier mindset at scale? I will definitely be reading into some of the CFAR resources, just curious to hear from you.
Certainly agree with your point about donating to Substacks / journalists. Could be very impactful to have a writeup of that somewhere here or on the EA forum.
I’m familiar with Galef’s ideas; I would place “soldiers” in category 2. But yes, the distinction is very subtle and I did not specify it well enough.
I believe that sufficiently well designed UI for navigating debates/arguments/discussions can make it very difficult for people to disguise soldier mindsets via obfuscated (intentional or unintentional) communication and reasoning.
Imagine, for example:
User creates a strongly worded post that features a clear strawman and/or blatantly skips over serious, in-depth prior discussion on the same topic.
An LLM categorizes the argument to properly situate it within prior discussion and notifies the user that they A) do not appear to have an accurate understanding of the original source—specifically pointing out why B) have not yet explored the X counterarguments coming after that line of reasoning, and the Y that come after that.
This could be seen as an enhanced version of “community notes” aimed at situating shallow, under-researched takes within a larger “map of human thought.”
Whether this can scale and outcompete current systems is unknown, but it does truly seem promising for the enhancement of public discourse and like a step in the right direction.
Appreciate the comment, was helpful in clarifying my thoughts.
Hmm, this does look interesting but I hadn’t really considered depth of 1-on-1 communication as a significant bottleneck. I also think the concept slightly falls apart when the users are not already knowledgeable / quick thinking / good at rigorous communication, as I’d guess there would be a steep learning curve.
Software wise, I think I’m aiming closer to better debates and debate tools, mainly because these things could be made public-facing and seem like an obvious use case for even present-day LLMs.
If you have any further thoughts regarding implementation of those things, I’d be eager to know. I’ll be trying to make my plans more specific.
even with perfect onboarding someone would still need to motivate themselves to read quite a large chunk of text to reach the pareto frontier.
Sure, I just don’t think this is a big issue because people bottlenecked by motivation probably won’t be serious contributors, and it’s worth improving things for those that will.
I think the main bottleneck isn’t tooling, it is how much time people are spending to make things legible for other people/newcomers.
Yep, I’m pretty much in agreement with your broader point here – most of this stuff fails at the user level. It seems like manually organizing content would be a good place to start for me.
At the same time, I want to explore how this can become a self-propelled process in the future. It seems like an obvious use case for LLMs, and presumably not a massive engineering task.
I appreciate the comments – hopefully you can weigh in as I continue looking into it.
I think Joe is arguing that unilateral capability restraint might exacerbate such risks.
For example, to the extent that safety-related considerations end up motivating or rationalizing especially drastic forms of international action aimed at shutting down or significantly restricting AI development in other countries, I think this could well be harmful even relative to more baseline forms of economic and military competition.
Like, the US government is all in on safety but China is not complying. Which would probably get intense, and rightfully so.
I still agree with you because:
I don’t think that is very likely at all; if one country/coalition is all-in on safety, I’d guess there is something motivating their fear that would apply to everyone else as well.
The default outcome of nobody being all-in on safety seems obviously much worse as you point out.
There is an argument to be made that great power conflict is justified in the credible expectation of total annihilation otherwise. Obviously this is extremely controversial as it can be easily misused, and IIRC has been used to deface the EA/rationalist community (or Nick Bostrom?) in the past, but I don’t think that takes away from it’s merit.
This previous LessWrong article seems extremely relevant and basically sketches out an example of the rough “strategic portfolio” for AI risk that you are arguing for.
In line with some of my recent posts, I’m starting to think there is a lot of value in:
Clearly defining consensus group strategy (among LW/EA/CG, for example) on “making the future go well.” This should include rough estimates from a variety of respected sources, a diverse portfolio of interventions, and explicitly communicated uncertainty / epistemic humility.
Designing info-UI tools to facilitate that process. Enabling effective deliberation, strategy adjustment, and maybe most importantly: easy-to-use interfaces for the general public. The goal being to convey community beliefs and disagreements in a very transparent and easy to understand way.
This is intentionally unspecific but I have outlined a couple particular ideas in previous posts and will continue to crystallize my suggestions / explain why I think this area has potential.
The AI futures model and related ecosystem is a great start but is limited to a handful of thinkers (Daniel and Eli) and a specific subset of information (forecasting timelines). Their work has already been quite impactful (read by JD Vance apparently) – why not work hard to apply and scale good information-interface-design to broader community strategy?
neo’s Shortform
What are some research directions for “improving coordination?”
In light of a recent post and comment, and several months of thinking, I have come to the position that one of our (humanity’s) biggest problems is that we suck at precise coordination at every level.
This is not very specifically defined but I am trying to gesture at a problem area I think is super important. Some thoughts to convey my intuition here:
If the extreme risk of the AI development trajectory is as true and obvious as many believe (everyone’s life at risk), humanity’s thinking about it should appear a lot more sophisticated than it does now.
For the last few years Eliezer has basically been throwing his hands up in exasperation at the incompetence of the world and many have shifted to public-facing communication, presumably believing that trying to convince AI insiders is hopeless.
Broadly, I think there are two cases of problems with coordination:
Two people/groups genuinely agree to honest, rigorous exchange of information, but can’t effectively coordinate.
Someone is withholding information or doesn’t really want to coordinate in the first place.
I think the first problem is workable, and if improved sufficiently, makes progress on the second problem by clearly exposing parties that are avoiding productive exchange.
Specifically, I think there is a lot of progress to be made with augmenting the exchange of information between people. I think LessWrong, the knowledge commons arguably at the frontier of ensuring humanity’s survival, is lacking in features for this purpose. Maybe because most users here are already conscientious and strongly value truth-seeking, which makes improvement seem less necessary.
Hopefully I’m making this line of thought clear enough. Key points:
Trustworthy, robust, and future-proof governance is the ultimate problem for humanity and anything else is a band-aid on a bullet hole.
Highly effective coordination is part of that problem, and better information exchange/clarity is a subset of that.
I think LessWrong can become an exceptionally effective prototype of this, and this could be very high leverage because of how proximate it is to the frontier of AI. Happy to expand more here.
I am interested in situating my thinking better here. Who is working on this sort of thing? I know TsviBT has explored improvements to debate, Richard Ngo / Samo Burja are exploring broader political manifestations, Forethought has published adjacent work. Is there anything I’m missing? Very interested in contributing here and think it’s a clear place where the ball is being dropped.
This hypothesis seems pretty compelling. I wonder what the strongest counterarguments are; i.e reasons that impeding prosaic alignment is a bad strategy for ensuring a good future?
I can think of the following:
Labs can easily transfer talent from capabilities to safety work, so we end up with the same pace of progress but instead it’s driven by less safety-minded / x-risk pilled people.
Impeding prosaic alignment successfully slows down AI progress overall, derailing anticipated productivity gains and leading to an eventual replacement of the current paradigm with something less safe.
I doubt (1) because I intuitively assume that the talent pool is sufficiently limited for this kind of transfer to be hard—though I could be wrong. I think (2) implies a successful slowdown which I think would be net-positive, though I’m a bit unsure about the downstream effects of stagnating industry revenue. Overall, I think these arguments are weak rationalizations.
What am I missing?