You allude to self-deceptive reasoning in the introduction; What’s your perspective on the extent to which prominent figures in AI safety (both now and historically) who have acted in ways you describe as power-seeking were doing so consciously vs subconsciously? This feels like a pretty load-bearing distinction, especially when it comes to thinking about how to intervene to make things better (conscious ⇒ get better at trusting the right people and put prominent figures under more scrutiny, unconscious ⇒ that + much more self-reflection amongst other things).
Jason R Brown
Great post, thank you very much for writing this.
I’d be interested in your thoughts on DCI, particularly any criticism or concerns you might have about it. I’m intending on drafting a longer post on its philosophy and theory of change, but the rough summary is that I see it as a potential way of paradigmatizing the science of generalisaton.
I’d also be interested in your thoughts on the work being done by Resolution, and their hopes around understanding LLMs.
Great post!
I (and others) have had some similar thoughts / intuitions. DCI is one concrete research agenda I’m really keen on pursuing that is motivated by this, and it may be of interest to others who enjoyed this post.
We did not, but I would be very interested to see their results!
Great post!
Another argument in favour of this is that RL is not very information dense, and will probably stay that way for quite a while. Constitutional training / SDF probably conveys quite a lot of information, and so even if RL is scaled up these efforts will probably be scaled up too, and still contribute most of the bits used to pinpoint the model’s parameters after pre-training. Also one way RL will be scaled up is longer horizon tasks, which will provide roughly similar total information to short horizon ones (at least for RLVR / outcome-based RL), despite having much higher compute costs. The relatively small amount of information contained in an RL reward signal means it’s probably most effectively utilised as a pointer to / steering towards existing cognition, rather than specifying new complex cognitive machinery.
J-Lens is typically degenerate in the first 1⁄3 layers
Well they’re not wrong on that… just maybe not right in the way they intended
This seems like another useful tool in the white-box toolkit, and as a non-mech-interp person, I appreciated the contextualisation in the paper of this method vs other such tools and the relevant trade-offs.
However, this post seems to oversell the method compared to the paper. The paper concedes that the logit lens captures much of the workspace-like structure of the J-lens, frames the method as complementary to existing approaches, and includes quantitative comparisons to other methods (though they are buried in appendices). None of this makes it into the post, which doesn’t mention the logit lens at all. That said, I don’t think this should undermine the workspace finding itself, and if anything makes it stronger. Overall though, reading the post alone you’d probably think the J-lens was a much bigger leap than seems to actually be the case.
I’m also not super convinced by the comparisons to global workspace theory. It seems that one would expect the J-lens (from its construction) to act as a sort of buffer for tokens that might need to be emitted soon. To be fair, the paper does provide lots of evidence for the workspace interpretation vs others, but some examples still leave me unconvinced. For example, in the data-fabrication example, the first set of outputs seem to mostly be the model trying to predict future words such as “realistic” or “improved” (which is in fact the next word). And yet it appears to me from context that the information to deduce the data might be fake is already present at this point (“I need to optimize the data values. Let me update...”). Only later does the J-lens output seem to reflect this. So overall it looks like some mix of potential-next-token buffering as well as something higher-order, and leaning hard on the workspace analogy seems premature (the paper seems better hedged in this regard).
I don’t want to be too negative, this work does seem really cool, and the paper presents it carefully. It’s just that the post does make this seem much more shocking and exciting than it maybe actually is.
Great post!
Some minor points:
We don’t expect already-released models to be trying really hard to figure out whether they’re in deployment or being evaluated, though, so this still isn’t a very strong approach.
Is there a reason why you can’t just prompt them to do this?
Humans perceive subjective time
I imagine a smart enough model could check system clocks, server clocks, or possibly even program logs to get a sense of how much time has passed in-between various actions. The evaluation environment could try and spoof these sources of time info, but the model might be smart enough to detect that. Overall I think it’s probably fine to assume that how much time is passing is available to the LLMs if they wanted to know it.
I want to get a sense of how much overlap there is Manifold, LessWrong, and the AI Safety community, since some things I’m trying / planning to try rely quite strongly on it. As such, I’d greatly appreciate if people could quickly answer this poll on Manifold: https://manifold.markets/Jasonb/poll-where-are-you-in-the-intersect (Btw, I know I could’ve not made this poll on Manifold, but doing so was deliberate.)
Thanks for the response, this all makes a lot of sense and is very interesting.
Edward and I’s agenda is taking a fairly agnostic stance (at least to begin with) on what sorts of internal structure AI cognition has, and mostly aims at exploring and evaluating what kinds of structures seem likely or are useful to reason via. Modelling meta-agency a certain way or modelling agents as having agentic shards might help our dev-cog models fit the data better (including better predictions for out-of-distribution training pipelines), then this provides good evidence that these are useful structures for reasoning about agency, but we are not committed to them. In general we are just trying to seek out the most useful structures, whatever they may be. On that note, we did have some pre-preliminary evidence a while back for some more shard-like cognitive models, but for now we’ve found very simple structures (with no modelled meta-agency) to be good enough. I imagine as we scale the agenda in what we try to predict and what dev-cog models work well, if a meta-agentic-shaped gap starts appearing we can then start to properly tackle it.
On your points on Vanessa Kosoy and Veedracs, I broadly agree, but I think perhaps you are overestimating the degree of possible co-ordination on AI, and underestimating the effects of incentive pressures, relative to other people working on AI Safety. If I’m wrong about this and the following arguments / points are not new to you I apologise. Essentially, there are economic pressures to build strong goal-oriented AI. Powerful AI would be able to create lots of wealth autonomously running companies, or trading stocks (and actually now has). Because of this, AI researchers or the AI safety community do not get to dictate what the nature of AI should be (without sufficient co-ordination that seems very unlikely). Instead, we have to try and design a safe approach for AI that will itself outcompete or prevent the sorts of AIs that might be built due to humans just following the incentives to build CEO-bot. I think this is mostly what Vanessa is gesturing at when she says we need to be worried about systems with bad goals. It’s not “well some AI safety people might make AIs with bad goals so we need to be robust to that”, in which the argument you’re making of “hey maybe this is a bad framing for AI safety people to make” makes sense, but instead it’s “some non-safety-minded people might make CEO-bot and so we need to be robust to that / win in worlds where that happens”. CEO-bot is one specific example, there are other forms this AI might take, but I think the general point of there existing incentive pressures to create goal-oriented AI systems by people who might not care or understand about AI safety is enough to motivate AI safety people not being able to completely ignore the space of AI minds that are somewhat goal-oriented. I think Evan Hubinger explains / illustrates this point really well in his somewhat old post on 11 proposals for safe AI, where he uses training and performance competitiveness as two of his four desiderata. Hope that helps!
Definitely agreed! Especially if you’re exploring weird generalisation things. I’m hoping to do a short write-up of that paper for LW in the next week or so.
Interesting post! I think I mostly resonate with the core claim here as I understand it that reasoning about utility maximising AIs might not be super useful for AI safety, especially in the near term. However, I do think that a roughly agentic/goal-driven-entities lens is very useful for thinking about how intelligent autonomous systems might go bad. Maybe it’s worth separating out broader agentic-like-things vs the more specific VNM-style utility maximisers. I think there’s a risk here of criticising the latter and using this as evidence against reasoning in a style that still leverages the former at all.
I’d be interested in your thoughts on Shard Theory, which Roger also mentions, and additionally Developmental Cognitive Interpretability, my own research stance having tried to internalise a broader view of agency that leans less on the VNM-style of agency.
That said, it does seem like there are some strong arguments for concerning ones self with narrower views of agency that restrict to goal-directed systems. I think Vanessa Kosoy’s comment on one of the articles you link is quite good, and I also quite like this post. I’m interested in your thoughts on these too.
I really want to write more LessWrong posts and I have a few ideas / things sketched out. I thought it might be fun to use Manifold to allow people to bet on how well they might do or whether I’ll get round to writing them: https://manifold.markets/Jasonb/how-interesting-are-my-different-id.
Interesting post!
To what extent do you think this being useful / important is correlated with the Natural Abstraction Hypothesis? This feels like the crux to me.
If some version of NAH is correct, then maybe desirable personas cluster around the natural form of goodness / alignment we desire, and so extrapolating from them will likely be very useful. It might even be the ways in which they don’t cluster around this might be correctable in some natural way that still makes personas a useful starting point.
However, if NAH doesn’t hold, or at least doesn’t hold between humans/personas and superintelligences, then it does seem like personas are much less useful and are very unlikely to meaningfully capture / guide ASI towards the target we want.
I think this is very relevant if you’ve not already seen it: https://arxiv.org/abs/2406.06560
Another relevant market predicting in what years CoT monitoring will not work: https://manifold.markets/Jasonb/in-what-years-will-cot-monitoring-f?r=SmFzb25i
UPDATE:
The projects have been chosen! They are:
Cameron Tice: Goal Crystallisation
Puria Radmard & Shi Feng: Exploring more meta-cognitive capabilities of LLMs
Lennie Wells: Model organisms resisting generalisation
These markets will be left locked until their individual metrics are resolvable, all other markets for the un-chosen projects will be resolved N/A.
Thank you to everyone who traded on these markets, and special thanks to those who provided feedback about the research projects and the futarchy experiment itself.
Ahh I see what you mean now, thank you for the clarification.
I agree that in general people trying to exploit and Goodhart LW karma would be bad, though I hope the experiment would not contribute to this. Here, post karma is only being used as a measure, not as a target. The mentors and mentees gain nothing beyond what any other person would normally gain by their research project resulting in a highly-upvoted LW post. Predicted future post karma is just being used optimise over research ideas, and the space of ideas itself is very small (in this experiment) and I doubt we’ll get any serious Goodharting by selection of them that are perhaps not very good research but likely to produce particularly mimetic LW posts (and even then this is part of the motivation of having several metrics, so that none get too specifically optimised for).
There is perhaps an argument that those who have predicted a post would get high karma might want to manipulate it up to make their prediction come true, but those who predicted it would be lower have the opposite incentive. Regardless of that, that kind of manipulation is I think quite strictly prohibited by both LW and Manifold guidelines, and anyone caught doing it in a serious way would likely be severely reprimanded. In the worst case, if any of the metrics are seriously and obviously manipulated in a way that cannot be rectified, the relevant markets will be resolved N/A, though I think this happening is extremely low probability.
All that said, I think it is important to think about what more suitable / better metrics would be, if research futarchy was to become more common. I can certainly imagine a world where widespread use of LW post karma as a proxy for research success could have negative impacts on the LW ecosytem, though I hope by then there will have been more development and testing of robust measures beyond our starting point (which, for the record, I think is somewhat robust already).
Thank you for the suggestion!
Excellent post! I quite like the philosophy of AI debate and what it’s useful for as outlined here. I’ve found when talking to people who haven’t worked on / thought about this, this is the part they mostly get wrong, and it makes it harder for them to understand what the actual pros and cons of debate might be. This will be a good resource to point them to in future.
That said, I found this sentence a bit jarring / confusing:
It felt like the post was mostly building towards / presenting AI debate as the main solution to reward specification / outer alignment. Thus, it would help us not have super-intelligent reward-hackers, or models that were trained towards optimising for a flawed instantiation of our values that could be lethal if optimised for by a super-intelligence even if mostly beneficial when optimised for by a ~human level intelligence. But EM is not really that? To me it seems like a much more specific and niche threat model, where the risk comes from the emergent / weird generalisation effects, rather than the simple fact you’ve built a model with the wrong goal. I know this is a bit nit-picky but I just wanted to clarify how you saw this. Is debate in your eyes mostly an EM defence, and if we solved EM through other means it would be substantially less valuable, or would it still be our (current) best bet for providing good supervision to align models (even in distribution) beyond human levels of intelligence?