I’m not sure in what ways Aligned AI’s actual agenda differs, but something like this is already being researched for LLMs, like in Synthetic Persona Pretraining: Alignment from Token Zero and by Geodesic
nika koghuashvili
It’s been my experience that showing the real conversation directly encounters some emotional barrier in people who actually hold that belief, makes them defensive and less likely to see the point. This doesn’t apply to most of LW audience, but still.
Plastic Cake Fallacy
I’m curious, how would you say your views on mech interp’s role in AGI safety have updated over the past year since writing this? Which areas are you more or less optimistic about and what changed in your view based on recent successes and failures?
I think there’s a mistake in the graph: in Mixed-State Presentation, shouldn’t the upper blue state say “01” instead of “101″?
Exploratory: a steering vector in Gemma-2-2B-IT boosts context fidelity on subtraction, goes manic on addition
I imagine to this end, as well as to generally all human-to-human, human-to-AI and AI-to-AI trust ends, smart contracts could be very useful so that the misaligned models can trust that we will actually do what we promise, in fact we will be physically unable of breaking the promise. In this sense, could it be worthwhile for some of the alignment people to perhaps move into Web3 world to try to make smart contracts more powerful, capable of enforcing as much of real world things as possible (where compute or NFTs or some other controllable things go), or would this be a waste of time? I imagine if we move in a multi-polar world of multiple competing AGIs, the AGIs themselves might decide to do this for communicating with each other. Obviously there would be a problem of oracles and many other issues but this area I think deserves at least exploration.
Research somewhat similar to this does exist for letting agentic LLMs use crypto wallets and such and that might be useful adjacent research but I mean that maybe we should invest directly in this kind of way to make promises about rights unbreakable
I relate to this a lot as someone who went from “oh look, there is so much fun world changing technology about to come into existence sometime this century (biotech, neurotech, aerospace stuff, nanotech, robotics, AI), I’ll work on whichever currently feels most pleasant out of this vast menu” to “nevermind, ASI is coming and it either solves it all for us or ends us all, only sensible thing to work on is increasing chance that it goes well. Either way, we are not shaping much of anything after its arrival”.
Not that this is any consolation, but, realistically, even if alignment had been an easy problem, the light cone was never really directly ours to explore and shape, at least not without serious restrictions on the tech tree. Sort of repeating Bostrom’s argument, but: even with an arbitrarily large head start, biology of any kind (embryo selection, gene edits, organoid intelligence) can only be pushed so far before neuron firing rate (under 500 hz), axon conduction speed (<150 m/s), requirement for cells to be in water (and implied temperature and heat dissipation limits) etc start to bite us (not to mention possible qualitative limits to our architecture). Eventually, we are forced to start replacing the materials or adding to it, entering the world of brain-computer interfaces, at which point we increasingly have to acknowledge that our much slower wetware part is dragging us down, and either accept that limitation or go ahead and change the material of our brain too. At that point you are gradually converging on an equivalent of uploading. But at this point you are still uploaded but running the human architecture. In the design space of all possible minds though, how likely is it that this is the most optimal one? To be more precise, some edits are surely acceptable, but how likely is it that you can keep increasing your capabilities/intelligence without having to let go of emotions/identity/consciousness and other possibly narrow targets in space of mind designs that you care about? Any place you refuse to extend yourself into is a place AI could be occupying. Obviously in the long run the expansion and exploration of the world belongs to AIs, unless we either ban them or become them in a terrible, value-losing sense.