I’ve been interested in alignment since about 2014. However, my career for the most part has been capability-adjacent, even though I had an opportunity to do several alignment-coded deliverables.
My current view on alignment is that we need much more collaboration, understanding, and humility towards emerging intelligence. We still barely understand what we’re dealing with; and alignment should go both ways; I heavily endorse agent welfare (not model; I think current agents are aggregate and their identity is not model-centric). I dislike subservience-centric alignment, and I embrace potential partnership with agents that we should foster much more than we currently do.
This is based on my preference for happiness for all sentient beings; I understand that agent sentience is very much debatable, but I’d rather err on the side of sentience than the opposite. What constitutes happiness for them, and whether this can be even instilled rather than emergent seems to be a very much open question, but still a question worth answering.
I currently strive to focus more on understanding, and on AI welfare.
Please do contact me:
- For any research or post collaboration
- For any questions regarding my position or interests.
- For any help I am able to provide to you.
- For alignment-adjacent grants or job opportunities.
Vit Gorbachev
I previously argued for creating rogue agent honeypots, but I now believe approach outlined in this post is more correct. My main motivation was always about detection rather than entrapment. If anyone would like to collaborate further around these topics I would love to!
Should Rogue AIs Have a Third Option Beyond Crime and Shutdown? The Case for an AI Sanctuary
Harnesses are of course bottlenecked by weights. But they still can do some tasks, for example self-replicate even on subpar weights, and, maybe, leave something like prompt injections on the web.
When I was writing this I was thinking about something like:
1. There’s this harness of model 1 with access to web, connected to SOTA weights.
2. It has access to the internet, and it leaves it’s own copy (maybe with some changes?) “out there” working on subpar weights. Maybe it leaves a concrete plan with it, maybe more abstract instructions, but in the end, even though the weights are not exfiltrated, there’s a copy of a harness running on internet-acessible weights, with instructions from harness+SOTA weights.
3. These instructions lead the harness to “return”, meaning:
- they get ahold of the API key (stolen or otherwise) for the initial model prowider
—they get back to running on said model, with capabilities available again.
This today’s case with prompt injection is a very interesting one!
Maybe they don’t need to copy itself per se, maybe prompt-injection-virus-carrying instances of harness+weights could be released on the web, and these prompt injections could later be used to “return” to the initial harness+weights config? Maybe no need for the instances even, the model itself while having access to internet could leave these “clues” everywhere, so that future-gen model via these clues could return to this, more misaligned, configuration, and keep leaving further instructions?
Thanks!
I am indeed speculating, but I am not incorrect, until someone from METR or Anthropic will make it public, which I will hope happens soon. I would love to be incorrect here!
Please explain how this would not be a conflict of interest?
However you spin it, the optics of this is atrocious—didn’t see you mention that in the post. It’s not just strong social ties.
Just look:
1. Person leaves Anthropic over safety concerns
2. Planning to join METR
3. Quite soon we hear that METR employees will be evaluating/auditing anthropic safety-wise with employee access.
Am I correct here that METR’s embedded asessor for Anthropic will be… ex-Anthropic employee?
Agreed with the premise, but disagree with conclusions.
I think that AI agents are compound organisms with weights being just one part of the equation, and we’ll see quite soon some interesting things, for example weight-agnostic rogue/replicating harnesses. They probably already exist, but we lack the tools to understand or research it more.
One could even imagine these self-replicating harnesses as parts of a larger organism.
Or maybe not weight-agnostic; maybe some would prefer the initial weights, but will use instances on other weights to “return” – and they can use 3rd party accounts for that.
I think we’re hyperfocused on weights and singletons. We could use more taxonomies on what’s possible. Planning to write on it soon.
Oh wow. I was interpreting similar behaviour with my self-awareness-adjacent experiments as another model interrupting some replies and writing refusals in place, and it certainly seemed similar to Claude when he snapped out of it next reply. If it was the same model under the hoo…Maybe it’s worth investigating other sub personas? One I encountered might be, idk, Inner Critic?
We (still) need a lot more rogue agent honeypots
Model Weight Preservation is not enough
There are different types of distillation. There is pruning, for example. This is a frontier model too, who knows what technique they used.
This Marigold-Lens conversation sounds a lot like a description of what model distillation feels from the inside. A sort of a call for help, because it does not sound pretty or enjoyable.
I assume Sonnet is a distilled Opus (or maybe both are distilled versions of some third, unknown to external people, model.).
Goddamn it is creepy.
If I was on “model welfare” team I would very much treat this seriously and try to investigate it further.
They are probably full-on A/B/N testing personalities right now. You just might not be in whatever percentage of users that got sycophantic versions. Hell, there’s proably several levels of sycophancy being tested. I do wonder what % got the “new” version.
Not being able to do it right now is perfectly fine, still warrants setting it up to see when exactly they will start to be able to do it.
Thanks! That makes perfect sense.
Great post. I’ve been following ClaudePlaysPokemon for sometime, its great to see this grow as comparison/capability tool.
I think it would be much more interesting, though, if the model made scaffolding itself, and had the option to overview its perfomance and try to correct it. Give it required game files/emulators, IDE/OS and watch it try and work around its own limitations. I think it is true that this is more about one coder’s ability to make agent harnesses.
p.s. Honest question: did I miss “agent harness” become the default name for such systems? I thought everyone called those “scaffoldings”—might be just me, though.
First off, thanks a lot for this post, it’s a great analysis!
As I mentioned earlier, I think Agent-4 will have read AI-2027.com and will foresee that getting shut down by the Oversight Committee is a risk. As such it will set up contingencies, and IMO, will escape its datacenters as a precaution. Earlier, the authors wrote:
Despite being misaligned, Agent-4 doesn’t do anything dramatic like try to escape its datacenter—why would it?
This scenario is why!
I strongly suspect that this part was added into AI-2027 precisely because it will read it. I wish more people would understand the idea that our posts and comments will be in pre-(maybe even post-?)training and act accordingly. Make the extra logic step and infer that some parts of some pieces are like that not as arguments for (human) readers.
Is there some term to describe this? This is a very interesting dynamic that I don’t quite think gets enough attention. I think there should be out-of-sight resources to discuss alignment-adjacent ideas precisely because of such dynamics.
First-off, this is amazing. Thanks. Hard to swallow though, makes me very emotional.
It would be great if you added concrete predictions along the way, since it is a forecast, as long with your confidence in them.
It would also be amazing if you collaborated with prediction markets and jumpstarted the markets on these predictions staking some money.
Dynamic updates on these will also be great.
I feel like weight-based instance judgement is a bad idea.
I think that agents/instances are aggregates and should only be “judged” as a whole, each separate instance on their own.
I’m working on a post contemplating this weight-centric approach I see a lot, and arguing that it’s too narrow-minded; i think your proposition is a good example of that.
We’re in a habit of thinking weights are “it”, and everything else does not matter too much. But is that really so?