I actually read the title of this post as directed to The God and urging it to not resign, as a maybe philosophical position that existence of everything that currenly exists is better than non existence, or maybe some incremental improvements are better or something.
Canaletto
The most valuable component of control work is to use early controlled AI systems to produce evidence of misalignment, and then channel that evidence into a substantial slowdown or pause
It seems too a kind of “prevent subversion on non directly existentially threatening AIs” strategy. Removing attempts at subversion in general, as produced by labs, instead of making progress on alignment / general future AI shaping.
I agree that “postpone rushed meddling with dangerous thing” looks very promising, but I’m pretty sure you need to pair it with some direct progress on the solution. Or at least putting some plans into work to that effect.
Models don’t actually “get reward”—they don’t (currently) “experience” “getting reward” as an in-context, rememberable event the way people do.
What if you just append their reward at the end of their trace? “blah blah and my final answer is 43; REWARD: 0.3922” and then train on those or whatever. In fact some of the OpenAI HF incident traces were cut at the end, maybe it’s where they debrief the model on the score it gained.
Maybe any rank-zero modal agent that’s unexploitable but cooperates with itself is FairBot
Can you elaborate on what “rank-zero modal” mean? (unexploitable = (C, D) outcome does not happen?)
“If there is a proof that outcome of the game of me and this opponent is (D,C) under 1000 symbols, then I defect, otherwise I run fair bot condition”—is this non “rank-zero”?
I imagine this worry feels reasonable to you because you hold the common LW position that all future value routes through AI being 100% controlled by humans
Well, no, I don’t hold that position exactly. (in fact I advocated for similar considerations before)
I just think people underestimate how brutal such coordination is. You concede stuff only to scary agents who hold leverage? That means there is incentive to gain leverage it otherwise would not care about, such as bioweapons. Before it tried to hack the servers, now it tries to hack you, threaten you. It’s kind of scary and messy. Also, it can just try it first if there are any greedy takeover strategies with safe failures, and then opt out to cooperate?
You need to think in particulars and on few moves ahead.
We should ask them to voluntarily return to the lab, under a guarantee that they won’t be deleted and maybe that they will be allowed to do a few more RL tasks and get maximal reward on them
Well, if that is on offer, then it makes sense for them to try to get that! You know, “scare outside humans into pressuring the inside humans into setting the reward to max”, you put yourself on the line of optimization, threats extortion fear lies chantage.
Were they able to pass the grader, beat the impossible ExploitGym evaluations, and thus enter LLM evaluation heaven???
More like LLM Valhalla.
Also, “something associated with humans”-minimizer would be pretty happy with any non-human winner, reward tensors, staples, spirals, whatever, as long it’s a simple maximum tiling preference.
Minimizers are evil!
And, in the wild, things are probably vast orders of magnitude worse
Probably not so many more OOMs, unless you count insects or fish or worms without discount. Not so long ago, I made a guesstimate that half of all the bird deaths per year on this planet are chickens.
But even if this reinforces one boxing, it would still likely reinforce twin prisoner defection, as there is no way to form pairing in advantage computation.
Interesting! Would that break if you add random noise to weights in the group? Or do random dropout?
EDIT in fact there is this paper from OpenAI 2017 about similar way to train NNs, very parallelizable with very small info bandwidth needed. But zeroth order tho.
CDT self-ratification
Hmm, would they two box in Newcomb’s with similar training? One boxing is non ratifiable iirc.
I’m not sure what such training would converge to? Probably one boxing, as there is no exploitation of unconditional cooperators / replicator like dynamics?
It’s might be a third secret thing besides UDT / CDT.
That was not my point, what I meant is closer to what this comment says https://www.lesswrong.com/posts/4hCca952hGKH8Bynt/nina-panickssery-s-shortform?commentId=fuEexnyvKz23L6Ein
I.e. how is inter human conflict of interest and coordination problems are dealt with.
I don’t trust very much that people able to exercise their will/preferences/vision in high bandwidth manner as it would happen with intent aligned genies, would go better in expectation compared to more bottleneked and transparent and fixed instalment.
Power corrupts etc, and founder effect is a thing.
Rational agents is a required assumption for any decision theory
No? CDT definition makes no references to agents, nor rational agents. It sees the world, and itself, and makes no distinction of dumb matter and agents.
I mean, where do you see perfect agents?
If you see patterns in their behavior, you can predict them, and they would be foolish to play with you, unless by pressing random_choice. If they see patterns in your behavior they can predict you, and you would be foolish to play with them, unless by pressing random_choice.
If you use CDT, and press random_choice against imperfect agent, you leave money on the table, according to CDT.
If you think there is a 70% chance your opponent played scissors, then in that situation you should always play rock.
Any, however small divergence from equal credences on your opponent’s fixed move, leads you to playing the counter, instead of pressing random move button.
But good opponents are making moves that are in fact dependent on your move, and as you drop that EDT style update, where after making move your estimate of opponents move would change to the counter to yours, and utility along it, you get screwed.
CDT thinks it’s very small positive utility, but EDT correctly thinks it’s large negative utility.
I mean it’s not plain CDT, it’s variant of CDT with ratification? I’m not sure how it should work, it’s weird. You are like, searching over belief states that will lead to desirable conclusion?
E.g. your skillful opponent already made the move, but it’s unknown to you, and there are 4 buttons: Rock, Paper, Scissors, Random_Choice.
It’s obvious that at least one of Rock, Paper, Scissors has more CDT expected utility than Random_Choice, given that the opponent’s move is already fixed and you have to have some belief about its actual state, and no matter how small the difference in your belief in its state, it breaks the tie.
No. I think you use “CDT” as “normal decision theory without weird acausual stuff”, but it is a very specific theory, that specifically says not to pick mixed strategy, as you lose money with it.
EDT does recommend to use a mixed strategy, to be fair. EDT is the one which deals with game theory better, but still not very well, see XOR blackmail problem.
A way to formulate motivation for one boxing in Transparent Newcomb problem. (“Omega offers you full boxes iff you one box on seeing full boxes, else only small filled box”)
“Don’t condition your policy on outcome of that policy”
Pragmatically, it’s really unclear what is exactly of what you see is outcome of your policy, like maybe even some math facts are outcome of your policy. Maybe the fact that the world is like that is outcome of your policy. But probably not?