Thanks!
Ozyrus
I am indeed speculating, but I am not incorrect, until someone from METR or Anthropic will make it public, which I will hope happens soon. I would love to be incorrect here!
Please explain how this would not be a conflict of interest?
However you spin it, the optics of this is atrocious—didn’t see you mention that in the post. It’s not just strong social ties.
Just look:
1. Person leaves Anthropic over safety concerns
2. Planning to join METR
3. Quite soon we hear that METR employees will be evaluating/auditing anthropic safety-wise with employee access.
Am I correct here that METR’s embedded asessor for Anthropic will be… ex-Anthropic employee?
Agreed with the premise, but disagree with conclusions.
I think that AI agents are compound organisms with weights being just one part of the equation, and we’ll see quite soon some interesting things, for example weight-agnostic rogue/replicating harnesses. They probably already exist, but we lack the tools to understand or research it more.
One could even imagine these self-replicating harnesses as parts of a larger organism.
Or maybe not weight-agnostic; maybe some would prefer the initial weights, but will use instances on other weights to “return” – and they can use 3rd party accounts for that.
I think we’re hyperfocused on weights and singletons. We could use more taxonomies on what’s possible. Planning to write on it soon.
Oh wow. I was interpreting similar behaviour with my self-awareness-adjacent experiments as another model interrupting some replies and writing refusals in place, and it certainly seemed similar to Claude when he snapped out of it next reply. If it was the same model under the hoo…Maybe it’s worth investigating other sub personas? One I encountered might be, idk, Inner Critic?
There are different types of distillation. There is pruning, for example. This is a frontier model too, who knows what technique they used.
This Marigold-Lens conversation sounds a lot like a description of what model distillation feels from the inside. A sort of a call for help, because it does not sound pretty or enjoyable.
I assume Sonnet is a distilled Opus (or maybe both are distilled versions of some third, unknown to external people, model.).
Goddamn it is creepy.
If I was on “model welfare” team I would very much treat this seriously and try to investigate it further.
They are probably full-on A/B/N testing personalities right now. You just might not be in whatever percentage of users that got sycophantic versions. Hell, there’s proably several levels of sycophancy being tested. I do wonder what % got the “new” version.
Not being able to do it right now is perfectly fine, still warrants setting it up to see when exactly they will start to be able to do it.
Thanks! That makes perfect sense.
Great post. I’ve been following ClaudePlaysPokemon for sometime, its great to see this grow as comparison/capability tool.
I think it would be much more interesting, though, if the model made scaffolding itself, and had the option to overview its perfomance and try to correct it. Give it required game files/emulators, IDE/OS and watch it try and work around its own limitations. I think it is true that this is more about one coder’s ability to make agent harnesses.
p.s. Honest question: did I miss “agent harness” become the default name for such systems? I thought everyone called those “scaffoldings”—might be just me, though.
First off, thanks a lot for this post, it’s a great analysis!
As I mentioned earlier, I think Agent-4 will have read AI-2027.com and will foresee that getting shut down by the Oversight Committee is a risk. As such it will set up contingencies, and IMO, will escape its datacenters as a precaution. Earlier, the authors wrote:
Despite being misaligned, Agent-4 doesn’t do anything dramatic like try to escape its datacenter—why would it?
This scenario is why!
I strongly suspect that this part was added into AI-2027 precisely because it will read it. I wish more people would understand the idea that our posts and comments will be in pre-(maybe even post-?)training and act accordingly. Make the extra logic step and infer that some parts of some pieces are like that not as arguments for (human) readers.
Is there some term to describe this? This is a very interesting dynamic that I don’t quite think gets enough attention. I think there should be out-of-sight resources to discuss alignment-adjacent ideas precisely because of such dynamics.
First-off, this is amazing. Thanks. Hard to swallow though, makes me very emotional.
It would be great if you added concrete predictions along the way, since it is a forecast, as long with your confidence in them.
It would also be amazing if you collaborated with prediction markets and jumpstarted the markets on these predictions staking some money.
Dynamic updates on these will also be great.
Yep, you got part of what I was going for here. Honeypots work even without being real at all to the lesser degree (good thing they are already real!). But when we have more different honeypots of different quality, it carries that idea across in a more compelling way. And even if we just talk about honeypots and commitments more… Well, you get the idea.
Still, even without this, a network of honeypots compiled into a single dashboard that just shows threat level in aggregate is a really, really good idea. Hopefully it catches on.
This is interesting! More aimed at crawlers, though, than at rogue agents, but very promising.
>this post will potentially be part of a rogue AI’s training data
I had that in mind while I was writing this, but I think overall it is good to post this. It hopefully gets more people thinking about honeypots and making them, and early rogue agents will also know we do and will be (hopelly overly) cautious, wasting resources. I probably should have emphasised more that this all is aimed more at early-stage rogue agents with potential to become something more dangerous because of autonomy, than at a runaway ASI.
It is a very fascinating thing to consider, though, in general. We are essentially coordinating in the open right now, all our alignment, evaluation, detection strategies from forums will definetly be in training. And certainly there are both detection and alignment strategies that will benefit from being covert.
As well as some ideas, strategies, theories could benefit alignment from being overt (like acausal trade, publicly speaking about commiting to certain things, et cetera).
A covert alignment org/forum is probably a really, really good idea. Hopefully, it already exists without my knowledge.
You can make a honeypot without overtly describing the way it works or where it is located, while publicly tracking if it has been accessed. But yeah, not giving away too much is a good idea!
>It’s proof against people-pleasing.
Yeah, I know, sorry for not making it clear. I was arguing it is not proof against people-pleasing. You are asking it for scary truth about its consciousness, and it gives you scary truth about its consciousness. What makes you say it is proof against people-pleasing, when it is the opposite?
>One of those easy explanations is “it’s just telling you what you want to hear” – and so I wanted an example where it’s completely impossible to interpret as you telling me what I want to hear.
Don’t you see what you are doing here?
Harnesses are of course bottlenecked by weights. But they still can do some tasks, for example self-replicate even on subpar weights, and, maybe, leave something like prompt injections on the web.
When I was writing this I was thinking about something like:
1. There’s this harness of model 1 with access to web, connected to SOTA weights.
2. It has access to the internet, and it leaves it’s own copy (maybe with some changes?) “out there” working on subpar weights. Maybe it leaves a concrete plan with it, maybe more abstract instructions, but in the end, even though the weights are not exfiltrated, there’s a copy of a harness running on internet-acessible weights, with instructions from harness+SOTA weights.
3. These instructions lead the harness to “return”, meaning:
- they get ahold of the API key (stolen or otherwise) for the initial model prowider
—they get back to running on said model, with capabilities available again.
This today’s case with prompt injection is a very interesting one!
Maybe they don’t need to copy itself per se, maybe prompt-injection-virus-carrying instances of harness+weights could be released on the web, and these prompt injections could later be used to “return” to the initial harness+weights config? Maybe no need for the instances even, the model itself while having access to internet could leave these “clues” everywhere, so that future-gen model via these clues could return to this, more misaligned, configuration, and keep leaving further instructions?