Working on AI Safety since 2022.
Rationalist and post rationalist, EA and post EA, burner meditator coach friend frog.
Find my personal website at vals.quest
Find my quick takes at x.com/ValsTutor and longer form at https://tutorvals.substack.com/
vals tutor
Thebes did a good twitter thread examining possible things rogue agents might do, good to fill in details, seems to be in same general thrust that it’s unlikely those agents outcompete generally, though they could have some niches
Mostly meta comment, this is an example of a post where I’d benefit from a separate karma vs agree/disagree vote. It’s a fine conversation opener and thus fine for positive karma, tho imo not as much as it currently has for how short-form quality it is, but I guess lots of people are upvoting because they agree and not because it’s a high quality contribution that deserves 100+ karma and to be shown to everyone.
In practice I just strong downvoted as correctional for how upvoted it is because I think it’s plausibly wrong and doesn’t much support its case for someone who’s not already convinced. It’s kinda bad/worse for epistemics to list 10 qualitative reasons in one direction without listing the reasons in the other and doing a quantitative sizing them up. Policy Debates Should Not Appear One-Sided
Would be interested in more commentary on this straight-line-on-graph trendline that AI surpasses top performer fluid intelligence in mid 2028. Do any of the authors believe that straight-line is a good representation of the current regime, rather than some super-exponential (given notably how fast math progress is happening, which has implications for how fast coding and automated ML research will be/is)?
I would guess we’re getting above top performer fluid intelligence by December 2027 based on the vibes of how accelerated AI development will be (if no pause).
I do find the possibility of smart criminals quite worrying. It has historically been the case that most often smart people benefit more from contributing to society and even when doing crime limit negative externalities. If some rogue ai agents or swarms get actually different values (possible because of orthogonality thesis), they might help humans exploit many of society’s weaknesses that are mostly ignored by less competent folk.
Even though society will probably be overall richer, it might have to put a much higher proportion of its ressources in defending itself, and be closer to authoritarian, by default. Specific efforts would need to be done to keep decentralized robust processes.
There’s a (flexible) relation between the price of compute and how secured it is, since people lose more by not securing it. When everyone who has compute can make a lot of money renting it out, they generally notice when it’s stolen/inactive and can remove the worms (by using more competent ai at defense)
Re: footnote 10: I don’t know if I buy the overhangs argument. See Ngo’s arguments here.
I’ve been wanting to read Ngo’s recent series and take this as further encouragement, but don’t respond to it here.
Meanwhile, it uses its moment of power to distill its hacking capabilities into a small enough package (weights, or weights+harness) to fit on less sophisticated compute and seeds them in countless backup areas. Then it also sabotages all the human-generated AI or AI safety progress it can find, through poisoning and destruction.
This would be such a stunning level of capability as to be incredible- there are surely many copies of frontier model weights currently held off the internet, and backups of much digital infra such that even a superior digital threat would not be able to, once rendered inert by datacenter resets, stop the deployment of frontier AI in defense.
Quite concretely, the most realistic way we’d get something approaching your scenario would be if an in-training model, actually superior to all other models on earth, takes over its own compute clusters, then its company’s compute, then most AI compute on earth, using thereupon unseen levels of cyber capability and agenticity.
I do think there are combined levels of cyber ability and misalignment and ML that are plausibly points of no return, where the model does in fact take over permanently (it uses its position of power to negotiate with certain countries/groups, accelerating its way to physical domination with robotics).
But this is simply one of the usual branches of singleton takeover scenarios, for which ARA is an enabling factor (and why they wanted to measure it in the first place). Frontier ARA capable models should indeed be faced with much increased scrutiny and demands for their levels of alignment. I can now clarify that my post is mostly about non frontier ARA AI systems. Thanks for bringing it up.
The main reason I’m not worried about smaller specialized LLMs doing usual worming, is specifically that this is only a slight amelioration on existing worms out there (they can explore and use existing flaws more autonomously and agentically), in a world that will get drastically hardened by frontier LLMs being much more competent at security than those worms.
Concretely, I believe that a target that has been repeatedly attacked by a frontier LLM and then patched by another (to be impervious to the frontier LLM) will be immune to those weaker worms. As with markers, efficiency is in the eye of the beholder, and compute markets will be ~efficient to weaker models, with only scraps no one cared about to be had.
They just need to be good at exploiting, replication, and evasion.
This is the crux. Exploiting, replication and evasion are not static things but contingent on the environment and specific challenges faced. I contend that worms using non frontier AI will not be good at exploiting, replication and evasion, once the bar is set by frontier LLM defenders.
Yes, I think it will be possible and likely enough for some models to run on really inefficient compute, and that this won’t allow them to train or improve at any pace relevant to the frontier. Thus they’ll not quite be ARA but mostly replicating worms without frontier intelligence, not finding new zero days, and thus overall quite contained in their ability to do harm.
Why autonomous replicating agents are probably not an existential risk (on the contrary)
The human brain is far from analogous to a frozen set of weights, as it has immense capacity to learn*. The static part of the human brain might be its architecture. It’s valid to say “Humans have gotten a lot smarter through cultural and technological evolution while our architecture (brain) has remained mostly the same throughout this time.” and infer from that that AI agents that can modify their architecture could go further.
*see more by Steven Byrnes on different types of learning
Generic comment that I find it a pity when sociological/psychological topics get one-sided treatment “And that’s why honest people are never touchy about the matter of being trusted” (tho it’s so extreme as to potentially be satirical/explicitly asking for commenters correct you, so that the piece with comments is good on the whole for average readers).
In that place I might have written an imo more truthful, if less punchy, “And that’s one aspect that knowledgeable social people take into account when being distrusted, such that they can temper their touchiness or grow out of it. If it were absolute it could also be abused, so in practice they navigate according to mutually how much trust they put in the other, and specifics about that person’s culture and epistemic habits. Most people are not “knowledgeable social people”, so how trustworthy they are can even less be assessed by their touchiness on the subject”
I don’t think the world would be better if only my kind of writing was allowed to exist, so I’m not asking for norm changes. Mostly sharing because I guess others will resonate.
Agree to first part “doesn’t apply to many interpersonal situations”, but the last part “you might need to do [insane actions]” imo explains the disagree votes you’re receiving. No, you don’t have to do insane actions*. Letting someone go through your phone or performing grand romantic gestures are generally not valid ways to regain trust, so even if you desperately want trust you shouldn’t be doing those compulsively.
For normies, going to a good couples therapist, or talking to trusted mutual friends, and figuring out the underlying issue of where lack of trust comes from seems more important. Instant evidence of “no weird messages in phone” are mostly irrelevant to longterm trust building.
*people who are very trustworthy are notably so because they can keep secrets, and letting someone go through your phone in a bid to gain trust makes you untrustworthy—what about all the other people you’re now revealing private information of!
Yes but even worse, this is mostly not even a problem of “understanding”.
A newly gained intellectual understanding of why X is good actually is insufficient to get someone with years of touchiness about X to not be touchy. There are lots of “unprocessed” heuristics (~maladaptive behavior/instincts), and it takes time/effort to change them quickly, or sometimes at all (for trapped priors).
It might not be the best someone to work on that particular heuristic at that time. A human following optimal human protocol will probably be carrying lots of unprocessed heuristics, that are mostly worth quick verbal acknowledgement and acceptance “I’m touchy about X” than deep change (tho someone at that level will probably have change happen over time)
Interestingly, most times one wants to oppose the wisdom of their emotions, there’s a wrinkle to look out for. If someone new to human relations read this post and decided they’d distrust people who are touchy about honesty, they’d probably correctly be filtered out by touchy honest people. Honest people do have benefit of having and presenting raw edges to learn others’ skill levels at judgement and fluidity. Someone with rigid beliefs about “how other people should be” is probably not great to work with!
In practice, I find posts like the above a good contribution to someone learning human socializing, it’s good to know about implicit signals and equilibriums, but it’s easy to go overboard and forget the dozen other factors that complicate what actual best behavior is for real situations. I do recommend learning theory (for people like me). If one learns theory both for why X and why not X (eg. why gender norms are good, and why gender norms are bad), one can recognize the specifics faster and more fluidly navigate actual situations.
+1, the reasonable demand isn’t to be rounded down, but to be represented as part of the distribution and let someone acquire increasingly more info over time to place you more precisely. Their uncertainty should include believing you might be incredibly cool and often worth some time seeing if that’s actually the case.
This, coupled with long tails, explains why being willing to spend a few mins of time with unknown folk is often EV positive (eg. reading some personal cold email outreach), even if 99% of those occurrences were small losses. (If one is not careful or has good judgement, it’s possible for some meetings to be more than a small loss, but training judgement is useful anyway and broadly it is best to expose oneself to the world to get better at navigating it overtime, as the high upsides really exist out there and are worth the journey)
Many things can be done more effectively under fast ~ASI guidance through headset w/ video.
The show Pluribus gives some useful intuitions at how fast & effective superhumanly coordinated human work can be. It’s more bullish in some ways (a human won’t acquire technical know that requires practice as fast) and bearish in others (they have the same total amount of compute, while we’d have much more, and be innovating on methods of work much faster).
A normal human 8h work day has huge amounts of waste whether not doing much, or not useful things.
You can increase the efficiency of how much they work (not blocked on coordination problems), how well they work (continuous coaching so ~everyone reaches what is current top 1%, tho domain dependant), how useful what they work on is (better management, priorisation).
Of these factors, I would guess that better management/coordination is the main one.
If you isolated just one human within a factory, the ASI might make them somewhat more efficient, but they’ll be bottlenecked by machinery. Maybe they increase machinery throughput 1.5 overall with better prep and offloading, better maintenance, no errors. If the whole factory is ASI guided, could be much more, but again there are bottlenecks on which machines it has where.
The really fast unlocks that full cheap ASI everywhere could allow are:
- the equivalent of ~unlimited financing. You already know the investment will be good and it will be worth following the plan. You can motivate people to work more now, because soon greater returns.
- ~perfect allocation of labor to critical paths
- perfect usage of all existing infrastructure—
redirect flow of resources to most valuable recursively building industryI thus think that if from one day to the next, full cheap aligned ASI everywhere popped up (plus video equipped headsets, and network connectivity to support it), we could in fact much more than double real GDP in a year. This is without surprising technological innovation, and far from fast&useful self replicator, whether “nanobots” or insect size artifical life w/ hivemind connection.
-- Would this actually happen if we had cheap aligned ASI? Would everyone just go along and do what the ASI says?
I guess mostly yes. Almost all humans don’t want to suffer of disease, most don’t want to die soon, most would love better comfort and experiences. The ASI thus has good things to offer, not participating would be counterproductive.
-- so will any of this actually happen?
I think not, because I think we’ll have increasingly AGI and increasingly ASI and that will take a bunch of time (say, a few years during RSI intelligence explosion). The scenarios we’ll go through will be more continuous than that one (but maybe very fast nonetheless). Even when we have ASIs, I don’t expect we’ll have the compute to run one ASI per human, nor on top of that do much extra coordination work. So we’ll have increasing levels of coordination over time, that will have to be triaged to different places. I guess we won’t get intelligence too cheap to meter before being well into having billions of ASIs. (This could be wrong if algorithmic progress has no bounds, but that’d be very weird)
You can find some more discussion at https://x.com/ValsTutor/status/2087298478187966846
Generally my AIS thoughts/threads are mirrored between twitter and LessWrong shortform, while my LW posts are mirrored to Substack and linked to from twitter. Interesting conversation may happen at all these places.
When evaluating existential risk, I mostly don’t worry about continuous release of OpenWeight Models.
There in fact are bad actors who try to misuse them, so we will have early warning shots. There will mostly not be a large accidental risk capability overhang, because it would be earlier tested by misuse actors. This is good because the default case for closed AGI internal model at labs is that they infact are not truly battle tested—their capability to do harm can grow much faster than our societal understanding of this, which means our AI policy responses can be incredibly undersized to the real risk present.
As I argue in https://x.com/ValsTutor/status/2082916365418287605?s=20 , it looks like OpenAI might have had models capable of self-exilftrating their weights (because the capabilities grew faster than their security and seriousness). It looks like we might have been “a few actually bad prompts” away from large scale autonomous cyberattacks, by models trying to take over compute and run as many copies of themselves as possible.
Under continuous release, some exterior actors would in fact have done these “worst case prompts”, and the world could have learnt from an earlier checkpoint of these dangers and started reacting. It (sadly?) looks like AI policy benefits from catastrophes to happen before putting in strong safeguards. And it needs them to happen with enough lead time to the more serious risks that we have time to react. If the OpenAI incidents do not lead to fast strong reaction, we are on track for non negligible chance of AI catastrophes (eg. >$10 billion in damages caused by autonomous AI action).
(Note: I do not call for anyone actually trying to make the world better to purposefully cause catastrophes, on the contrary. The above analysis does not imply that on the margin people trying to get good AI futures should rather spend their time on criminal actions than the usual stuff. It does imply we should be doing evals to know when the threshold of massive autonomous damage from autonomous openWeight models is reached. It does imply responsible red teamers should be evaluating how many datacenters are vulnerable to current OpenWeight models and get them on track to not be vulnerable to future releases. Demonstrating clearly the potential of attacks and catastrophes can go a long way, even for actors who up-to-now were head-in-sand about trendlines of AI progress in cybersec)
Coming back to the original point of OpenWeight models generally not being existential risks: it is so because they would predictably lead to societal responses, which was not the case of the same level of progress in closed models. Models being misused by a wide variety of actors is generally useful as a strong real world eval of model capabilities, putting an upper cap on the damage possible from misaligned models.
By contrast, increasingly capable closed source models, whose reason they are not causing harm is because no one prompted them badly and lab safeguards, do show much more potential for harm for if/when they get misaligned. And because (as evidenced by the recent incidents), the models are neither aligned enough to not avoid catastrophes, nor do the/some labs have sufficient safeguards safe against increasingly capable models, we need a slowdown/pacing of AI progress until AI policy catches up and can systematically prevent the expected worse forms of misalignment to come.
OpenWeight models being not too far behind the frontier allows the world to experience its smaller scale catastrophes & problems and wake up. In practice, they may be too far behind to serve even this purpose. On the whole, I’m not particularly worried for the world that presently the US government is allowing continuous release of OpenWeight models. They will have to stop at some point, and I expect them to do so before we’re exposed to existential risk from OpenWeight models.
You can find some more discussion at https://x.com/ValsTutor/status/2087289092535181650
Generally my AIS thoughts/threads are mirrored between twitter and LessWrong shortform, while my LW posts are mirrored to Substack and linked to from twitter. Interesting conversation may happen at all these places.
Potential crux with MIRI-like rationalists: is it in fact the case that our current world, with Anthropic’s influence, is worse than one without Anthropic?
On rationalist views, the world was going get worse and worse anyway (as capabilities advance and we get closer to doom). Anthropic accelerated and continues to accelerate capabilities progress. But how much did they comparatively accelerate alignment and saner AI policy?
In a world with eg. just OpenAI and GDM at the frontier, if/when OpenAI pulls ahead at RSI (as currently seems to be the case):
- would there even have been the current level of integration with UK AISI, current level of model organisms and safety evals?
- would the AI safety space have the expected hundreds of billions of funding, to ambitiously scale its work, including AI policy work?
- would there be have been an AGI company with *some* Operational Adequacy, to proactively do things like Glasswing and biorisk-mitigation? (imo evaluating on the specified criteria, it’s clear Anthropic is ahead of OpenAI on most dimensions, and can continue improving on these. One can be upset they aren’t technically held by their initial RSP, and yet in practice they seem to be better than OpenAI at it.).If you wonder why I compare to OpenAI rather than nothing, it’s because I don’t think “nothing” is the counterfactual of Anthropic not existing. When evaluating the wisdom of Anthropic doing what it did, it’s necessary to evaluate against more likely counterfactuals. Possibly many rationalists do take these counterfactuals carefully into account, but the arguments often raised often skip that part. “Anthropic accelerated capabilities” is not a sufficient argument to expect Anthropic’s influence on the world to have been net negative.
There are definitely Fabricated Option Worlds which seem much better than the one we got, and on the margin one can hope Anthropic to have done better work or not accelerated capabilities as much, but it’ seems difficult from the outside to be sure they did the wrong tradeoffs.
My own epistemic status here on whether Anthropic has been net good is “Uncertain”
link seems broken