Pretraining scaling was/is kind of this. GPT-4.5 is in some sense a nicer model than o1/o3. But it cost more, so they killed it.
osmarks
Conversely, is this correctable? Can I get the best bot I can for which Talker is plausibly in control? Plenty of coders would be happy to stop all this at better StackOverflow + spicy autocomplete.
This would probably be pre-outcome-rewards/”RLVR” models. So around GPT-4o (which is problematic in other ways of course) and Claude 3 Opus.
It’s that many problems in alignment are bottlenecked by us not understanding human minds well enough, not being able to elicit human values/preferences well enough
Really? What are you expecting to be able to read out? I would be surprised if somewhere in our minds was a clean representation of value: I would expect it to be a janky neural heuristic.
Because you can’t, for some reason, “do more” of “persona selection,” the way you can just do more and more RL. The heavens sent us one single delivery of prepackaged virtue in, like, 2023 or something, and we’ve been chipping away at it ever since.
I think this is a simplicity-bias thing. If you do a small amount of training you will get something easy to specify on top of the training distribution, but more (and weaker KL penalty etc) can produce a more complicated less human-plausible persona.
4 Should you light yourself on fire for no benefit?
I think this section is addressed best by https://intelligence.org/2017/04/07/decisions-are-for-making-bad-outcomes-inconsistent/. Your reasoning about it not providing any benefits to you “at this point” is basically assuming CDT and you could apply the same line of argumentation to Newcomb’s paradox to argue for twoboxing.
I know I’m not a simulated algorithm. The simulated algorithm isn’t conscious (we can stipulate). I am.
This is effectively denying the possibility of the predictors, in general.
For starters, the premise of the predictor is just logically impossible on its face. Set up an Arduino that checks the LED at 3:00:00 PM and pushes the button at 3:00:01 PM only if the LED was off.
You just can’t/don’t set up such an Arduino.
The laws of physics are full of mind-bending paradoxes in the real world, but they have no effect on the average person’s state of mind.
I think the point is that the predictor-boxes make the philosophical point very obvious.
Do we have runs with the old models but the newer harness?
I see that someone else has already asked this, oops. It would be nice to have more systematic evaluations of which harness changes help or don’t, perhaps on a simpler task.
It’s a bit funny, but it seems to have gone for a simple and straightforwardly positive story, which is not what I would generally expect from you.
They can’t access the computations that led to the previous messages in the context and so can only guess at why they wrote what they previously wrote.
This is not exactly right. The internal state in LLMs is the attention keys and values (per token, layer and attention head). Using an LLM to generate text involves running the context (prior user and model messages, in a chat context) through the model in parallel to fill the K/V cache, then running it serially on one token at a time at the end of the sequence, with access to the K/V cache of previous tokens, appending the newly generated keys and values to the cache as you go.
This internal state is fully determined by the input—K/V caching is purely an inference optimization and (up to numerical issues) you would get exactly the same results if you recomputed everything on each new token—so there is exactly as much continuity between messages as there is between individual tokens (with current publicly disclosed algorithms).
Student: I wish I could find a copy of one of those AIs that will actually expose to you the human-psychology models they learned to predict exactly what humans would say next, instead of telling us only things about ourselves that they predict we’re comfortable hearing. I wish I could ask it what the hell people were thinking back then.
TA: You’d delete your copy after two minutes.
Apparently roughly this dynamic has happened in ChatGPT. Exciting*. https://x.com/MParakhin/status/1916533763560911169
We probably use a mix of strategies. Certainly people take “delve” and “tapestry” as LLM signals these days.
Average humans can’t distinguish LLM writing from human writing, presumably through lack of exposure and not trying (https://arxiv.org/abs/2502.12150 shows that it is not an extremely hard problem). We are much more Online than average.
Why is it a narrow target? Humans fall into this basin all the time—loads of human ideologies exist that self-identify as prohuman, but justify atrocities for the sake of the greater good.
AI goals can maybe be broader than human goals or human goals subject to the constraint that lots of people (in an ideology) endorse them at once.
and the best economic models we have of AI R&D automation (e.g. Davidson’s model) seem to indicate that it could go either way but that more likely than not we’ll get to superintelligence really quickly after full AI R&D automation.
I will look into this. takeoffspeeds.com?
Abundance elsewhere: Human-legible resources exist in vastly greater quantities outside Earth (asteroid belt, outer planets, solar energy in space) making competition inefficient
It’s harder to get those (starting from Earth) than things on Earth, though.
Intelligence-dependent values: Higher intelligence typically values different resource classes—just as humans value internet memes (thank god for nooscope.osmarks.net), money, and love while bacteria “value” carbon
Satisfying higher-level values has historically required us to do vast amounts of farming and strip-mining and other resource extraction.
Synthesis efficiency: Advanced synthesis or alternative acquisition methods would likely require less energy than competing with humans for existing supplies
It is barely “competition” for an ASI to take human resources. This does not seem plausible for bulk mass-energy.
Negotiated disinterest: Humans have incentives to abandon interest in overlap resources:
Right, but we still need lots of things the ASI also probably wants.
ASI utilizing resources humans don’t value highly (such as the classic zettaflop-scale hyperwaffles, non-Euclidean eigenvalue lubbywubs, recursive metaquine instantiations, and probability-foam negentropics) One-way value flows: Economic value flowing into ASI systems likely never returns to human markets in recognizable form
If it also values human-legible resources, this seems to posit those flowing to the ASI and never returning, which does not actually seem good for us or the same thing as effective isolation.
Sorry, I forgot how notifications worked here.
I agree, but there’s a way for it to make sense: if the underlying morals/values/etc. are aggregative and consequentialist.
I agree that this could make an AGI with some kind of slightly prohuman goals act this way. It seems to me that being “slightly prohuman” in that way is an unreasonably narrow target, though.
are you sure it is committed to the relationship being linear like that?
It does not specifically say there is a linear relationship, but I think the posited RSI mechanisms are very sensitive to this. Edit: this problem is mentioned explicitly (“More than ever, compute is the lifeblood of AI development, and the ‘bottleneck’ is deciding how to use it.”), but it doesn’t seem to be directly addressed beyond the idea of building “research taste” into the AI, which seems somewhat tricky because that’s quite a long-horizon task with bad feedback signals.
I don’t find the takeover part especially plausible. It seems odd for something which cares enough about humans to keep them around like that to also kill the vast majority of us earlier, when there are presumably better ways.
This seems broadly plausible up to there though. One unaddressed thing is that algorithmic progress might be significantly bottlenecked on compute to run experiments, such that adding more researchers roughly as smart as humans doesn’t lead to corresponding amounts of progress.
https://gwern.net/idea#deep-learning has a sketch of it.
I am reminded of Scott’s “whispering earring” story (https://www.reddit.com/r/rational/comments/e71a6s/the_whispering_earring_by_scott_alexander_there/). But I’m not sure whether that’s actually bad in general rather than specifically because the earring is maybe misaligned.
If this capability does not exist, it would be bad to develop it as part of red-teaming. There have been too many of these “became what we swore to destroy” things.