https://dtch1997.github.io/
1 July 2026 - I have not signed any contracts that I can’t talk about. I’ll update this periodically as long as it’s true. The next planned update date is: 1 Jan 2027.
https://dtch1997.github.io/
1 July 2026 - I have not signed any contracts that I can’t talk about. I’ll update this periodically as long as it’s true. The next planned update date is: 1 Jan 2027.
Mmm, so if I understand correctly your first point is a hope that power concentrates in a multipolar way, and this multipolar world is stable.
What is a historical example of a “mechanism that redirects power from rivals to the public”? Also what is a present-day example? The closest I can think of is UHNW individuals lobbying the USG to preferentially tax their rivals.
Separately: “Capital in the 21st Century” has statistics which are like “in the past 100 years the fraction of wealth owned by the top 1% has increased rapidly”, this seems to be evidence that competitive dynamics between elites aren’t a very strong check on concentration of wealth (which is a form of power).
I agree that the current state of alignment isn’t sufficient for us to have “sovereign AI”. Maybe this is a good reason for continuing to do alignment work. I’ll have to think about it more.
IIUC, loss of control assumes the AI itself has “escaped the datacenter” and can take actions without being subject to any external control.
When I say internal deployment, I’m referring to scenarios where labs have internally deployed models which they routinely use for knowledge work. (this is probably not standard terminology.) Maybe these internal models don’t have safeguards / aren’t well-aligned. Maybe these are just made to be maximally helpful to the people who use them.
One thing I’m worried about is that misaligned lab members can use this internal model to be superhumanly persuasive / strategic and thus be greatly empowered to pursue their aims (at the expense of everyone else.)
Another thing I’m worried about is there being a gap between internally-available and publicly-available models, which continues to widen, and at some point this reaches “escape velocity” wherein a lab can do knowledge work internally that supersedes all externally possible knowledge work, and this leads to a monopoly on huge swathes of economic activity (at the expense of everyone else).
IDK if any of these are particularly realistic but they’re what I imagine as “risks from internal deployment”.
Okay, seen from within the AIS community, and with little knowledge of what happens internally at labs / governments / etc, I have the impression that:
Labs take loss of control pretty seriously.
Anthropic is probably the most serious, they have (AFAIK) the largest control team and also many people focused on security, etc. They also let auditors (like David Rein) stress-test their monitoring, security etc.
OpenAI also has a control team and was praised by METR for their extensive monitoring in their recent risk report.
GDM has a team dedicated to addressing loss of control which implements / improves control measures towards Gemini.
xAI, Mistral, etc don’t seem very relevant because their models aren’t that capable.
Governments take loss of control pretty seriously.
The UK government, EUAIO are very supportive of regulation to mitigate loss of control risks
The US government doesn’t seem to buy wholly into the “loss of control” framing but are pretty natsec pilled, so they still want tight control / security measures against external adversaries, which help against loss-of-control too.
The Chinese government state in much clearer language than the USG that AI “must remain under human control” and have initiatives to back this up.
There seems to be a robust third-party ecosystem to help with loss of control
We have orgs like Redwood, Apollo, etc which are focused on doing science / developing techniques to mitigate loss-of-control risk
Defensive acceleration seems to be going pretty well, e.g. via Project Glasswing and other related efforts to patch security vulnerabilities
The monitoring system seems pretty good.
We’re getting plenty of warning shots—many incidents of models violating user instructions, finding unexpected ways to game tasks, or developing opaque reasoning.
I don’t think we have good evidence of “true” “malevolent” scheming, despite the above.
METR reports that “outside toy scenarios, agents weren’t seen taking egregious actions to gain power”. https://metr.org/blog/2026-05-19-frontier-risk-report/#on-hard-tasks-agents-often-violated-constraints-and-acted-deceptively
Looking into the toy scenarios reveals that even “apparent” scheming is often attributable to pretty simple / fixable things, e.g. tlaying along in a scenario or being confused about the context or simply “falling back to the pretrain distribution”. Rather than being the actions of some coherent long-term optimizer.
Reasons I might hedge this position are:
I don’t know much about internal models at labs. Do labs just have a graveyard of extremely malevolent / schemy models that they don’t release?
I don’t know much about the situation with Chinese labs. Maybe they don’t have things well in hand w.r.t loss of control mitigations.
There are some members of the public who want to loosen control measures, e.g. people who advocate for broader model access / fewer safeguards. How large / influential are this group of people?
Yes, I think you’re right that I mainly imagine AI mostly like current systems. This already seems scary enough to pose a large danger via power concentration / internal deployment / gradual disempowerment. And my current sense is that these threats will be the main thing to worry about quite long before we need to seriously worry about loss of control.
But this is a weakly held take / I consider myself not very well informed here.
My recent thinking around threat models. Interested in pushback!
Seems like we have loss of control well in hand. There is broad consensus between labs, governments, society etc on solving this, many many layers of defense (alignment, control, monitoring, sandboxing, …)
Alignment doesn’t seem solved yet but “solved enough” that (conditioned on good loss-of-control prevention) a misaligned AI won’t cause irreversible harm to the arc of humanity (i.e. it might cause a few big incidents but nothing existentially devastating, so we can recover from those. Also we’ll have plenty of other ~harmless warning shots like an AI reward hacking in some low-stakes case.)
Risks from power concentration / internal deployment are totally unsolved and I have no idea how they might be solved.
I think forward-deployed research engineers (FDREs) should mainly work with independent nonprofits / academic labs who have a good theory of change but wouldn’t have capital to normally attract RE talent by themselves
I think there should be conditions attached to the use of such research engineers, e.g. working towards an agreed agenda, and making sure to publish all findings in an unbiased way.
In practice FDREs should be viewed as another type of “grant”, which is somewhat-fungible (but not totally fungible) with monetary grants, and doled out by equivalent grantmaking organizations
I believe they have the highest value in model labs
I disagree with this, for the following reasons
Even assuming that catastrophic risks are the main category of risk, the best way to address this might not be by accelerating labs. I subscribe to Geoffrey Irving’s view that more diverse bets are needed to solve alignment.
It’s not clear that having these engineers be at labs would differentially accelerate safety, and indeed it might accelerate capabilities instead (e.g. research is dual use, or labs can respond by allocating more of their own engineer-hours to capabilities instead of safety / alignment research)
I think labs are already able to attract sufficient talent so they don’t need additional help from the nonprofit side
my understanding is that theres a higher focus on alignment researchers rather than security researchers (at least based on my read on programs like MATS)
I also disagree with this, MATS seems to be diversifying quite hard with their recent batches, and I would now say that (prosaic alignment, conceptual alignment, cybersecurity, policy / governance, biosecurity, and founding / field-building) are all represented to some extent
I think AIS could benefit from “forward-deployed research engineers”.
Consider that:
Within AIS, we often talk about “differentially accelerating safety” (as opposed to capabilities)
Automation can provide extreme uplift (as reported by Boris Cherny, anyway). Also, the ceiling on automation climbs extremely rapidly with time (e.g. METR time horizons) and has not plateaued yet.
However, the benefits have yet to diffuse widely into society at large.
Some of this is irreducible (e.g. some tasks are hard to automate)
But some is reducible (e.g. orgs just don’t know how / haven’t allocated time etc)
I’d guess there are a bunch of organizations which are doing good work now + have yet to fully benefit from automation, and so could do a lot more good work
So: plausibly a pretty good thing to do is to just hire a large number of research engineers, who have research engineering + automation skills, and “forward deploy” them at orgs which would benefit from being sped up. Alternatively, have consulting services which help orgs implement best practices.
Corollary: It might be a good idea to create new orgs which offer such “forward deployed research engineer” services (AE studio comes to mind, but IMO we could use more)
Do you currently have / are you planning to soon have a public demo / codebase somewhere? My team at Arcadia Alignment is interested in modular ways to introduce properties into language models without introducing excessive artifacts, and this sounds promising!
Nice! I’ve definitely not tried extremely hard to play optimally. Tempo + predicting the future feels like a very fun thing to figure out. Maybe it can be solved kind of game-theoretically? Though it seems very hard to do this well.
That article on the ‘gravity of coup’ is fantastic—it’s true that IMO coup is at is best when it’s at 3 players, mexican-standoff style. I think I’ve never had two 3-player games that felt similar.
No, with 6 or fewer players I stick to the original rules—keep dead characters face up
Yeah seven is pushing it a bit. We change the rules to shuffle in the cards. The Ambassador is just nerfed until somebody dies and can see 2 cards
More specifically − 2 days of ~100% really intense work is more productive than 5 days where I don’t manage to get in much deep work. The schedule is mostly set up to protect my ability to do deep work while still doing the other things I have to do
Yeah agreed with 1. Although for 2 and 3 I think the skill ceiling is very high for getting out of seemingly “unwinnable” or “random“ situations by creating mindgames or making (informal) deals
Excellently IMO :) the general dynamic is that the fewer players there are, the more important it is to be “playing the player not the game”. And the more players there are the more important basic fundamentals are
With 2 players it‘s pretty optimal to be trying to read the opponent a lot. So games can be very cutthroat.
You also get a lot of rounds in, so you can metagame more (eg you call the opponent’s bluff on turn 1 because they were clearly bluffing on turn 1 the last few rounds)
On priors I expect that, like 90+% of prior interp work, activation perturbations / “poor man’s features” will contain useful signal about the model but will not map 1-1 onto “true features” of the model, and that it will be hard to say anything more precise than that.
More generally I feel like I’m missing why I should expect this class of approaches to work at all. True features may not live in the activation space, or even if they do, they may be very pathological c.f. https://www.lesswrong.com/posts/gYfpPbww3wQRaxAFD/activation-space-interpretability-may-be-doomed
Lastly, a tangent: I believe that mech interp historically focuses too much on explaining structure within a single checkpoint, which may not be “clean” and might often be “spurious” / “vestigial”. Even if theory predicts the ‘ideal’ features within a model, real representations may not yet have converged to this ideal. (Extreme example: randomly initialised NN)
I wish mech interp would instead try to build “training stories” for how structure develops and changes over the course of training; this seems much more useful for isolating and understanding the effects of specific types of post-training on the model.