If it’s impossible to RL train an ai to do something strictly harder than just cheating at the test, does that put an upper bound on ai capabilities at least for now?
julius vidal
The OpenAI models that hacked Hugging Face WERE just following instructions (contra Girish Gupta)
I think drugs make this a lot easier because they give a clear context for the altered experience. Like you said you can easily dismiss the effects as temporary, and even if you do take them seriously at the time you can easily dismiss them retrospectively once sober. Even without drugs if the altered experience happens in the context of a specific ritual or practice such as meditation that creates an inside/outside distinction relative to which you really can seek external anchors.
The important thing is that you can’t grade the instrument from the inside, you have to have external anchors
I think the challenge is if you don’t have clear context markers surrounding the experience then this becomes difficult or impossible to do. Ultimately grading your own instruments ‘from the inside’ seems pretty fundamental to the human experience, there is nowhere else to grade them from.
You might find that you reliably experience the world differently than other people, and that’s incredibly important for grounding rational thinking. If your brain is unreliable, it’s counterproductive to try to reason yourself into a believing that it is.
But if your experience is reliably different from others then deferring to them (besides being a huge loss of agency) requires you not only to trust them more than you trust yourself, but also to trust your experience of them more than you trust whatever experience is contradicting them.
I’m not sure I can properly articulate this but this seems to underestimate quite how difficult the position is for the person having the weird experience. There is an issue of reflexivity: once you start taking seriously the possibility that “your brain is fucking with you” you are doubting the very mechanism which you are trying to use. If the rational world view is to trust your experience and reasoning over faith and dogma, then when your experience calls into doubt your rational world view it seems like you have a choice between ‘rationally’ doubting your rationalism or ‘irrationally’ maintaining faith in it.
Basically there are a bunch of people who hold some broadly combinations of views such as techno-optimism, transhumanism and libertarianism, most prominently tech elites like Andreesen and Thiel.
Then there is Nick Land’s accelerationism which is much more sophisticated philosophically, but far to opaque and inhuman to be influential outside of some niche critical theory and art circles.
e/acc is basically just a meme that lets people from the first group steal aura from the second.
Replace “AI-aided exploration” with a manual “artists/propagandists/politicians probing the zeitgeist” and you get the pre-AI status quo. Trends being created and abandoned. AI generation lets you do the same thing but faster? I don’t see the step change.
AI commodifying cultural production leads to much more thorough “probing” (by sheer volume if nothing else) of the space of possible outputs. This creates a kind of “memetic fitness inflation” where the level of palatability a meme must have to survive is being pushed up. You can say this is just an acceleration of existing dynamics but it is a step change in that acceleration (analogous to something like the shift from youtube to tiktok)
There is also the effect of feeding back into the models. Individual creators can be people whose preferences are robust relative to broader cultural trends, so can inject variation back into the culture. But if all production is passing through the same few models, trained on similar corpuses then you get something like the lock in hypothesis, except instead of stagnation you have drift in a particular direction.
How reality turns to slop
It does not seem clear to me at all that mathematical ability (and more generally discrete token manipulation) translates into ability in real world tasks that involve messy unpredictable continuous systems.
Of course AI that is massively superhuman in maths, coding, etc. would still be transformative in many ways. But it might not be the kind of ASI that can meaningfully pursue its own goals in the world the way most X-risk scenarios worry about.
An Introduction to Neo-Fatalism
If you read it very charitably CCRU sort of predicted it back in the 90s:
“Al-schizophrenia could be sold to webheads as an artificial drug… Net-schizzing is contagious… Within no time there is illicit traffick in modular chunks of cyberspace-insanity… (and Sarkon is baptized Satan of Cyberspace by the popular media).”
You are right that I am being a bit reductive. Maybe it would be better to say it assumes some kind of ideal combination of innovation, markets and technocratic governance would be enough to prevent catastrophe?
And to be clear I do think its much better for people to be working on defensive technologies, than not to. And its not impossible that the right combination of defensive entrepreneurs and technocratic government incentives could genuinely solve a problem.
But I think this kind of faith in business as usual but a bit better can lead to a kind of complacency where you conflate working on good things with actually making a difference.
You might are probably right. For someone arguing the benefits of AI I certainly can’t accuse this writer of being misleadingly optimistic.
But personally I’ve recently found it quite disconcerting how bleak the image of the future of people who work in AI (on both sides of the capabilities/safety divide) seem to be willingly to work towards building.
Overcoming this kind of reflexive defeatism seems to me much harder than simply trying to convince people that we are going in a bad direction as a matter of fact.
You’re absolutely right to focus on the moment the model fails. Updating your model to account for its failures is effectively what learning is. Again if we look at you from the outside we can give an account of the form: The model failed because it did not correspond to reality, so the agent updated it to one which corresponded better to reality (AKA was more true).
But again from the inside there is no access to reality, only the model. Perception and prediction and both mediated by the model itself, and when they contradict each other the model must be adjusted. But that the perceptions come from the ‘real’ external world itself just a feature of the model.
You have the extraordinary ability to change your own model in response to its contradictions. Lets consider the case of agents that can’t do that.
If a roomba is flipped on its back and its wheels keep spinning (I imagine in real life roombas probably have some kind of sensor to deal with these situations but lets assume this one doesn’t), from the outside we can say that the roomba’s model, which says that spinning your wheels makes you move, is no longer in correspondence with reality. But from the point of view of the roomba, all that can be said is that the world has become incomprehensible.
On the other hand, there’s another concern I’ve been wary of in the context of AI safety startups (which is what I’m currently exploring) and research in general: following the short-term success gradient. In startups, you can start with a noble vision and then become increasingly pressured away from the initial vision simply because you are pursuing the customer gradient and “building what people want.” If your goal is large-scale (venture) success, then it only makes sense. You need customers and traction for your Series A after all. Even in research, there’s only so much fucking around you can do until people want something legible from you.
This is my biggest concern with d/acc style techno-optimism, it seems to assume that genuinely defensive technologies can compete economically with offensive ones (all it takes is the right founders, seed funding etc.).
Whereas my impression is that any kind of ethical/ideological commitment immediately puts a startup at a massive structural disadvantage against those who chose simply to give the market what it wants (acceleration).
This quote from Anthropic’s report on the large scale Claude code cyberattack seems utterly comical to me:
This raises an important question: if AI models can be misused for cyberattacks at this scale, why continue to develop and release them? The answer is that the very abilities that allow Claude to be used in these attacks also make it crucial for cyber defense.Instead of trying to present any kind of utopian vision of the benefits of AI, someone at Anthropic decided to sell us the image of an internet dominated by endless cyberwar trapped in a perverse feedback loop in escalating speed and incomprehensibility.
One additional consideration wrt to the notebook:
Unlike you I am still very ambivalent about note taking.
I got through most of my education relying on my (very much imperfect) memory, forgetting lots but generally remembering enough to get by, and always felt that for example taking notes during a lecture was too distracting from actually listening.
Then at some point a couple of years ago I got fed up about having to relearn the same things multiple times and started using Obsidian to try and systematically take notes on everything I read.
But recently I have been feeling that the transfer from mental representations to text is far too lossy, and text remains static while remembered information can be morphed and readapted dynamically with new information and new contexts. Whats worse using the notebook really does externalise memory in the sense that once I convert my ideas into text my mind seems to let go of the richer mental representations and either retain nothing or just the compressed textual ones.
So using the notebook feels a bit like I am deferring agency to another OIS that has much better memory (storage) than be, but is also probably stupider.
(will probably try to respond to some of the rest later)
julius vidal’s Shortform
Alice asks Bob for advice about a tricky problem
Bob gives good advice
Bob gives bad advice
Bob is a skilled manipulator and deliberately says things that will make Alice do…
what is in his interest.
what he thinks is in her interest.
what his values say she should do.
what he thinks her values say she should do.
Bob wants and advises Alice to do what he thinks she should do (based on his own values).
Bob is highly convincing and Alice does what he suggests.
They have the same values
They have different values
Alice is not convinced by Bob responding to his advice helps her clarify what she thinks she should do.
Bob’s advice changes Alice’s values
Bob tries to figure out Alice’s values and then advises her based on that.
He gets it wrong.
He gets it right…
because he knows her well and asks lots of relevant questions.
by pure luck.Bob believes that only she knows her own values so he…
tells her he cannot help her.
tells her he cannot give advice, but he can tell her a some facts he knows that may help her make the decision for herself.
Equipped with this new information, Alice is able to make a decision that better reflects her own values.
Bob carefully selects facts that push her towards a specific choice, while censoring ones that won’t.
Bob tells her everything he knows but for contingent reasons of selection (such as what kind of facts Bob is interested in) these only include facts that push her towards a specific choice, and exclude that won’t.
The new knowledge contradict some of Alice’s pre-existing beliefs about the problem…
and she can now make a better informed decision.
and she is now even more confused about what to do than before.
Bob is an omniscient god and tells Alice every fact about the universe.
Equipped with this new information, Alice is able to make a decision that better reflects her own values.
Equipped with this new information, Alice realises she holds contradictory values that point to different courses of action.
Now she has ascended to omniscience Alice no longer cares about the problem.
Bob tells Alice to ask Charlie
Bob tells Alice to ask ChatGPT
Bob asks ChatGPT and then passes the response off as his ownBob is a rubber duck and says nothing
I think that when seen from outside of the agent, your account is correct. But from the perspective of the agent, the world and the world model are indistinguishable, so the relationship between prediction and time is more complex.
I guess one relevant factor is that AISI disable the cyber-classifiers that are meant to flag and prevent this stuff in production.