AI safety concerned since 2015, now an independent researcher. Currently working in AI Control.
deepy
I’m of the opinion that we need new agent harnesses that have all these modern features while being easy to apply control protocols to and run our evals on, without there being any ways for the model to hide what it’s doing due to some structural issue with the harness. Control is clearly a bit of an afterthought for Claude Code and Codex, even if their built-in control mechanisms are working well right now.
Also, do you have any examples of factored cognition being used by a misbehaving/red team agent successfully? I am looking into studying factored cognition as a control protocol with the expectation that it will make the monitor’s job easier, even when the orchestrating agent and the subagents are untrusted.
(Also, good post, and I broadly agree with everything)
Why the interest in a Dutch translation? Virtually every well-educated native Dutch speaker is also able to speak and read English very well and (low confidence in the following suggestion) there might even be a broad preference for the original English for non-fiction works on technical topics amongst native Dutch speakers.
Is it just for completeness since the publisher is going to produce a Dutch translation of the book regardless?
Future Whey, sold in Australia by Bulk Nutrients. One small correction: it also has extra BCAAs mixed in (the idea being a large surplus of BCAAs signals muscle protein synthesis more strongly)
You can get a generic EAA powder in other markets, it’s just not clear to the typical consumer that it counts as a protein powder.
Hello, u/Adhiraj and I independently did some research and the only good source of methionine we found is brazil nuts, which of course you probably don’t want to eat much of since they can give you selenium poisoning.
The conventional wisdom is to just eat more protein to compensate for the poor amino acid balance, but you’re going to be eating a lot, so I prefer to supplement methionine.
I have adding methionine in meals I cook and adding it to the plant based protein powder I use, precisely measuring how much I add. You can get individual amino acid supplements from iHerb and some larger health food shops.
Some sources of vegan protein go much much farther doing this (e.g. red lentils). I tend to eat soy TVP which goes somewhat further with methionine supplementation.
I also have a pure EAA protein powder which is entirely fermented and have a balanced amino acid profile. These are generally very poorly marketed supplements but they are widely available.
There’s no need to add it to your food rather than mix it in a drink, other than you probably want to get all your amino acids around the same time, and it helps if you’re doing meal prep and you don’t need to work out how much to supplement every time you go to eat.
There is apparently research suggesting that too much methionine can increase the risk of type 1 diabetes. I’m not going to go into that, but it makes sense to err on the side of not over-supplementing methionine.
We have been use this tool to help work out the amount of methionine we need to supplement, but it’s not a pleasant user experience: https://tools.myfooddata.com/protein-calculator
This is just my quick answer from memory, I may do a full writeup later.
Some thoughts as someone who has been eating plant-based for the past year and who thinks about the ethics constantly for fun:
I know someone on a mostly-beef diet who will probably develop health issues if they stay on a diet with plant-based food other than fruit for too long.
Beef seems to be less bad from an animal welfare perspective than most meats, at least in farms in Australia. I would probably still pay a premium for extra-ethical beef if I stopped eating . Dairy is probably bad everywhere except India.
Kangaroo meat seems like a clear-cut example of a “vegan meat” in practice. The Australian government sets quotas for culling them (they are overpopulated in deforested areas), very precise sharpshooters shoot them in the head, and some of the meat from culled kangaroos is sold. Demand for kangaroo meat has no effect on quotas. There is no counterfactual animal harm from buying kangaroo meat. Kangaroo suffering (and negative externalities in general) resulting from culling is very low at any rate, save for the occasional accidental killing of a mother kangaroo. I’ve had kangaroo meat maybe five or six times since going vegan, usually either to deal with the odd craving, to signal to family that I’m not a purist, or because someone had cooked it or a restaurant had it on a menu.
Avoiding honey/oysters seems like it has less impact than it’s worth for the time I’ve spent pondering the subject.
You’re right that wild-caught fish are not necessarily humanely treated. The ikejime process for killing fish is very humane compared to the cheaper and more suffocation method, but I’ve never seen fish advertised as ikejime (and they are probably an order of magnitude more expensive).
Social pressures in either direction are a thing. The time I spend with different groups of friends has changed since I went vegan, even though nobody has really objected to it. My flatmate reduced their meat intake significantly almost immediately after I went vegan, but they still eat meat when going out to restaurants with others. People seem to care about food a lot. People also seem to take cues from those around them.
If eating vegan had noticeably worsened my health or my concentration, I would have stopped, because I can probably do more net good in the world without those problems even if I had to eat the flesh of sentient creatures. Eating vegan and spending your time on other things are not orthogonal for everyone.
If you rely on evolutionary heuristics for diet you’re probably not going to be vegan, especially since we know B12 is a thing that’s difficult to get without animal products, and there could be other things you aren’t getting. I’m not really worried about this since I don’t see many short term issues for myself, the medium-term issues seem to be more reliably mitigated against now, and if anything I’ve heard that it helps longevity.
Vegans trying to gain muscle for whatever reason aren’t necessarily going to have a hard time as long as they remember to eat protein at all. Vegans trying to cut (lose fat without losing too much muscle) are probably going to not have quite the same success if they were an omnivore or otherwise have a very boring time with the food they eat (unless they eat kangaroo).
I seem to need to eat more to actually have enough energy if I’m not eating eggs/meat. It’s more expensive if you don’t like cooking, but the extra fibre is usually a good thing.
I’d say that not eating animal products seems like the correct choice if you can pull it off without too much trouble. It probably makes more and more sense the older you get. If you can afford to change your diet but not exclude animal products entirely, consider eating kangaroo if your country doesn’t ban its import. Consider eating kangaroo anyway.
There’s an analogy being drawn between the power of a hypothetical advanced alien civilization and the power of a superintelligent AI. If you agree that the hypothetical AI would be more powerful, and that an alien civilization capable of travelling to Earth would be a threat, then it follows that superintelligent AI is a threat.
I think most people here are in agreement that AI poses a huge risk, but are differ on how likely it is that we’re all going to die. A 20% chance we’re all going to die is very much worth trying to mitigate sensibly, and the OP says still it’s worth trying to mitigate a 99.9999% chance of human extinction in a similarly level-headed manner (even if the mechanics of doing the work are slightly different at that point).
Given how long it took me to conclude whether these were Eliezer’s true thoughts or a representation of his predicted thoughts in a somewhat probable future, I’m not sure whether I’d use the label “candid” to describe the post, at least without qualification.
While the post does contain a genuinely useful way of framing near-hopeless situations and a nuanced and relatively terse lesson in practical ethics, I would describe the post as an extremely next-level play in terms of its broader purpose (and leave it at that).
I think I’m more motivated by the thought that I am going to die soon, any children I might have in the future will die soon, my family, my friends, and their children are going to die soon, and any QALYs I think I’m buying are around 40% as valuable as I thought, more than undoing the income tax deduction I get for them.
It seems like wrangling my ADHD brain into looking for way to prevent catastrophe could be more worthwhile than working a high-paid job I can currently hyper-focus on (and probably more virtuous, too), unless I find that the probability of success is literally 0% despite what I think I know about Bayesian reasoning, in which case I’ll probably go into art or something.
Some AI control and software security relevant ideas that have been floating around in my head or in my research for the past 7 months or so:
AI agent environments should have a well-specified action space such that actions can be constrained, and environment state should be easy to represent and restore, such that “agentic monitors” can get full visibility into the environment state before approving actions. Right now most agent environments are nowhere near this.
AI agent task decomposition is not just good for factored cognition based control protocols, but also because in advance of a task, permissions can be scoped to match precisely what is needed for that task. Trusted models can be used to effectively scale the principle of least privilege to large swarms of agents.
For background: Factored cognition based control protocols (a) make it harder to come up with and coordinate an attack if multiple subagents and (b) means there are fewer actions that are “on task”, so there are fewer relevant actions that can be used to hide malicious effects from a monitor
Recently, LLMs seem capable of doing more bookkeeping work regarding program correctness and safety, so we should be aiming them at programming languages that enforce certain classes of correctness and safety in exchange for this now-cheap bookkeeping. The problem is that Haskell is the only relatively popular programming language that provides very legible guarantees about side-effects, for example. Languages like Austral (which is a well thought-out but abandoned project by a friend of mine) provide excellent guarantees, and with formal verification, programs in a well-specified language can be constrained arbitrarily.
Side note: Recent LLMs are also much more capable of doing formal verification work, and we should be aiming to do more of this. The bottleneck is human verification of high-level specifications or specifications of desirable properties of programs. Verification can improve program security tremendously, but your proof is only as good as your assumptions. Vulnerabilities have been found in formally-verified software before, because they had weak axioms (e.g. made assumptions about interrupts that weren’t true), so automated vulnerability scanning is still useful to us.
In response to the recently reported security incidents during RL training at large labs: AI control strategies used during RL training should be “asynchronous”, insofar as the monitor cannot affect the outcome of a trajectory. They can, however, halt training at a point where the model discovers a vulnerability or an impossible problem, before trajectories are produced that might reinforce bad behavior. The response here is not to downweight these trajectories, as this creates much selection pressure to avoid monitor detection and would effectively be training on the CoT. Instead, the response should be to resolve the issue with the RL environment: patching a vulnerability, or removing/correcting the impossible task.
In a typical control protocol, there may be certain kinds of actions that produce many false positives from the monitor. One strategy would be to train more precision in the monitor, but another is to provide a designated time period where those actions are allowed to occur such that they do minimal harm. The only example I have of this is “recon” operations where an agent explores a system or codebase. You could have a protocol where the untrusted agent is given its time to explore the codebase, but then must write out a message describing what it found that the monitor approves or paraphrases, then another agent (probably the same model) is given that message and is not allowed to perform any actions the monitor might deem “recon”. This seems like it might have usefulness costs, but it also seems like in this specific case it might inhibit scheming as well. Not that interested in this any more.
Capability systems (see https://en.wikipedia.org/wiki/Capability-based_security) seem appropriate for agentic environments. Monotonically attenuating capabilities in which you may only do one thing or another, but not both, might provide some enforcement of desirable security properties without relying on the monitor to do that.
It would be good to establish in general which models we can trust, and what we can trust them with exactly. I find my current workflows to rely on the models I use quite a bit. I have no deep justification for this being OK, and I don’t think this is a good status quo in a world where research sabotage is considered a possibility.
It seems insane to me just how bad computer security is, and how so few people are focused on improving it with the new models we have available. Reasons for this might be:
Tragedy of the commons since most of this work would not be profitable.
Resistance to moving away from the status quo (e.g. I don’t believe most operating systems are going to survive in their present form, and I have no idea what is going to happen with web browsers)
Limited access to sufficiently powerful models without safeguards outside OpenAI and Anthropic
It’s clear to me that you cannot get aligned models if you rely on RL and don’t have airtight security during training.
You can get a long way with this simply by not having any services that are connected to the internet in your RL environment. In the OpenAI/huggingface incident situation, simply do not have an NPM proxy and instead have a local copy of some relevant subset of NPM pages. This isn’t literal “air-gapping” but it removes all of the standout weakest links. It is much harder to escape both the gVisor process and KVM (or whichever hypervisor you use) than to find a bug in a proxy service.
However, I would still be concerned about bugs in whichever non-internet-connected services you have being exploited for coordination between episodes/agents. However, if you design these services right, it should be easy enough to prevent or mitigate this, since these services are supposed to be read only.
It is adversarial of me to say this, but improving security in the short term may increase x-risk since OpenAI’s proposed pause seems only to be until they sufficiently improve security, and not until there is a more robust story around alignment. We are reaching scary levels of capabilities and have had the blessing of multiple relatively benign warning shots. As a first pass, moving all capabilities research talent into interpretability research seems sane in a well-coordinated world.