If I wanted to be certain I was getting serial depth, I’d like, compose a bunch of hash functions together and ask the AI to evaluate the result on a bunch of inputs or something. But I mostly expect AIs to just do really badly at that, so I’m not sure it’d actually make a very good eval.
AprilSR
The obvious thing to me to try would be italicizing the host
you must not reveal that you think that
I think the actual demand is somewhat narrower than this, probably? Or at least, my instinct in a situation like this is often to be careful about the manner in which I discuss my thoughts, rather than to outright conceal them.
Someone could plausibly object even to the narrower requirement about the manner in which one reveals that, but I do think it’s narrower.
they may learn the lesson that the problem is difficult instead of I suck at the problem
Ah, that makes sense. I wasn’t quite able to piece together the dynamic that clearly from just the term, but this was enough for it to click.
I think the respectful approach is probably to just not capitalize the pronouns? Lots of Bibles don’t capitalize them!
Also, regarding corporate sin — the Catholic Church absolutely has a concept of “structures of sin” and similar such things, so I don’t think the idea presents any serious theological issues. Though maybe evangelicals have different theology?
I would generally describe myself as preferring FDT or LessWrong style decision theories generally, but I think in this particular setting they probably ought to recommend defection. (Even so, the result that Kimi moves CDT-wards makes sense, because it’s a more straightforward argument that CDT would defect in this case, and also there’s probably just a general association between CDT and taking actions described as prisoner’s dilemma defections.)
...I guess I was merely in high school when GPT-2 came out, but it sure seems pretty recent to me...
Yeah, I’m not totally sure but it does seem reasonably plausible to me that it requires caring about structural details that game theory tends to factor out. One might hope you could derive the relevant concept of “fairness” from the game theory, but… I’m not particularly sure you can.
I think this gets at one of the central problems with the concept of game-theoretic threats. I don’t know if anyone ever figured out an especially solid formal solution — it sort of comes down to finding some reasonable conception of a privileged null (or default) action, as I understand it.
I feel like I would struggle to get into a situation where I would regret hearing that the world wasn’t ending? Like, the world’s not ending! That’s fantastic!
I suppose I can try to draw a distinction between expecting to feel like “oh man why did I spend all those resources that was dumb in retrospect” instead of “we got lucky but I didn’t know that at the time, those were reasonable decisions”.
I think an acknowledgement that I have been wronged is actually often in and of itself something I care about a decent amount, even if it’s not associated with a change in behavior? To speculate, maybe it’s just like, reassuring to be reminded that someone cares about not harming me even if they’re not agreeing to make that their top priority or anything.
If in fact I’m not being given a sufficient amount of social capital to make up for the harm to me then yeah eventually it does start seeming fake. But I’m really confused about the idea that this would have to in all cases come through in specifically changing the behavior?
I think this is basically correct, but the English word “belief” is often considered to include opinions (because most people don’t draw an incredibly clear distinction between questions of fact and questions of value), so—like, you just have to be careful you’re drawing that distinction correctly, I suppose.
Oh, yeah, I absolutely agree.
A further note on logical correlations: I feel that people with a low level of familiarity with functional and evidential decision theory often over-estimate the extent of how many people they are logically correlated with here on Earth.
I don’t know if I would characterize it as an “over-estimate” personally. Mostly I think no one has an especially complete idea of how logical correlation actually works. I once asked Eliezer, and his tentative stab at it was
It ought to look similar to—though it is not defined as! -- the evidential update you’d probably perform on learning your own voting decision, if you’d never in your life seen any polls or gotten any info at all about how many people vote or for who; but you had even more information than now about which other voters are otherwise similar or dissimilar to you.
But I haven’t worked through this well enough to be very satisfied with my understanding.
(But of course I agree that thinking of FDT as sort of like magic mind control is probably (?) a little silly.)
I think making it marginally less convenient for authoritarian governments to catch dissidents could in practice be a pretty large benefit to them? It’s not clear to me how often authoritarian governments will even actually attempt the jail breaking, and if they do it probably matters just how difficult of a jail break it requires.
policy debates should not appear one-sided
your strength as a rationalist lies in your ability to argue for a belief being constrained by whether that belief is actually true
Hmm.
I like allowing it because it makes me happy when I see how downvoted they get.
I think different parts of a human can trade with each other, although I think they’re usually more abstract parts than the left vs right hemisphere.
I find 1 and 3 more intuitively plausible than 2?
I think someone trying that isn’t a bad idea, but Eliezer might not have the stamina. We should figure out who would make a good cowboy to be the face of AI safety to rural America.