When trying to get LLMs to imitate yourself or others, do you just prompt them “act like X for the duration of this conversation” or is there more to it than that? (I am thinking of things like putting your own or someone else’s corpus in context). I’ve occasionally wanted to hot-swap Claude’s personality with something more relatable. I haven’t had much luck, but also haven’t tried particularly hard.
Error
I’m not sure why I ran across this now and not when it was posted, but...that’s creepy, and very very good. Reminds me of Altered Carbon. I kept looking for a second layer of hidden meaning in the first poem, like you see in some SCPs, but didn’t find one.
...I don’t have anything actually useful to say, I’m commenting mostly because I know it’s a great dopamine hit to see people still discovering an existing story.
maybe you should try introducing it as the Irretrievability Problem rather than “oneshotness”
I’d like to suggest “Ironman Mode” (or whatever its best-known synonym is) as possibly memetically useful here. It refers to a difficulty modifier in certain games that prevents the equivalent of saying “oops” and restoring to 1929. Mistakes become permament, at least for that playthrough. The term isn’t a perfect match, because you can try again, but only by starting from scratch, nethack-style.
(“roguelike” was once a similar concept, but has been badly diluted in recent years)
One reason I think it might be a useful metaphor is that most players fear playing in ironman mode. Yes, it’s “only a game”, but progression takes time and effort, and thus losing it is a cost-in-reality—a cost that is intuitively real, in the sense that System 1 understands it, sees it as a real thing that can really happen, and that there can be no appeal to justice or mercy, and so shies away from it.
Another is that the few players who do play in such modes may be more likely to understand the concept you’re driving at. If you, say, have a Long War policy of treating a 99% hit rate as if it were 100%, you will fail, because sooner or later you will miss that shot in a situation where missing is fatal, and no amount of “but that should have fucking worked!” will save you. That doesn’t mean you never take such risks—but it does mean you develop a habit of explicitly checking that you can survive them first, and you learn a kind of principle: Do not YOLO if YOLO.
An amusing pair of XCOM moments: Long ago, having persuaded my then-partner to try EU, the first sectoid in the first real battle hits her first soldier through full cover for a critical hit that kills him. I later calculated the odds of that happening as ~2%.
Years later, showing off the same game to my current partner for the first time, the first thing she witnessed was my last soldier of the round taking a point-blank shot with a 98% hit rate. He missed, and the target immediately fatally shot him in the face.
That’s XCOM, baby!
Hence the “certain” qualifier. As the next comment down notes, EU/EW is honest (on Classic difficulty or above). Its popular Long War mod is honest unless requested otherwise. FFT is honest. I’m not sure what other games out there are honest but I’m sure those aren’t the only ones.
I’ve heard that XCOM2 lies flagrantly, hence I never bothered playing it. I’ve more recently heard that it’s honest on the highest difficulty only, but don’t know if that’s true.
Darkest Dungeon is almost honest, but not quite. I suddenly wonder if there’s a mod to make it actually honest, it doesn’t seem like it would take much.
Tactical games that display probabilities taught me more about probability and statistics than actual stats classes. If one wants an intuitive understanding of what numbers like 80%, or 20%, or 50% actually mean—instead of rounding them off to “yes, no, maybe”—then certain XCOM entries will teach that very quickly.
Along with anger-management skills.
(There is the hiccup where you need to identify a game and sometimes a difficulty level that doesn’t lie to you about its probabilities. Most do.)
If you like both, maybe make whichever one you don’t put in the Spotify album available as a single.
(not that it’s that important, since v1.1 is still available on Suno either way; it mostly just rubs at me because Suno’s app is awful.)
Possible bug report: Neither the Spotify nor Youtube releases seem to have lyrics for the songs.
I notice that I greatly prefer the “v1.1 (cover)” version from the LessOnline album to the Spotify one, but take that with salt; I’m not how much of the reaction is substantive and how much is just “after having it in regular rotation for ~1yr any change would sound wrong”.
Can’t believe no one has tried this yet. I miss the old site. This is what came up when I asked for LessWrong Classic; it’s not actually all that close except for the colorscheme, though.
The main thing I’d expect to get out of source vs. decompilation is the comments. Decompilation can tell you what the code does, but not what the author was thinking when they wrote it.
The Claude Code Source Leak
Bubblewrap has an
--unshare-netoption. I don’t use it—I’m not concerned about leaking code as long as it’s only my own code—but I would if I were working on anything sensitive.
Probably true now, but was it less true in the 1960s? It would be hard to replicate the Milgram experiment today, I think, even if its results were entirely accurate. Today, Milgram and similar experiments are well-known, and I’d expect an elevated level of paranoia among subjects that any seemingly-dramatic study may be deceptive. But those experiments were created and run in an environment that didn’t know of them, and I’d intuitively expect less paranoia and more trust in the experimenter.
The first way I bottleneck AI is by reviewing its requests for permissions to do stuff. Many people resolve this by YOLO mode, where AI can do everything it wants. I like my photos and production databases, so I don’t feel comfortable doing this. I am also worried about prompt injections from the web.
I’m pretty sure the right way to handle this is to run it in YOLO mode but inside a container or VM that can’t reach said photos and DBs. [1] I use a shell script that runs it under bubblewrap with a minimal set of
--bindoptions and a slightly wider set of--ro-bindoptions. I’d never previously used bubblewrap, but Claude was able to figure it out for me [2] . It did take several attempts to get it right.
I’m nowhere near Singapore, but if you’re taking suggestions, I’ve been wondering about the status of the Robbers Cave experiment.
I’ve expected something like this ever since LLMs grew web search limbs. I’m surprised the success rate is only 9⁄125, though if I understand things correctly that’s a lower bound.
I avoid using my legal name outside professional contexts, and I prefer un-googleable handles, because I don’t want HR departments digging through the personal side of my life and I don’t want randos doxxing me. I’m pretty sure a normie can’t connect my persona to my person. I think even a motivated techie would have at least some trouble. But I wouldn’t expect my precautions to hold up against, say, a twitter mob that decided it didn’t like my face. Given enough eyes, someone would find a connection.
LLMs give J. Random Nosy Bastard the eyes of a twitter mob. I expect that 9⁄125 rate to climb quickly, and I’m not sure what to do about that.
Woot!
...well, mostly woot. Finances might torpedo my plans to go. But I do plan to go if at all possible.
These two hit close to home.
I write frustratingly slowly. For anything topical, by the time I have something I’m happy with, the moment feels long gone. I’m still sitting on this year’s 90%-finished LO post in part because it feels embarrassing to put it up months later. My fanfiction tends to get written a decade after the source material left the conversation. And specifically here, with AI moving as fast as it has recently, the “moment” feels like it disappears almost before I even see the news.
And as for getting too old...I’m in my 40s and feel like I’m getting dumber every year. And can’t reassure myself with an IQ test, because the scoring compares against others the same age when what I want is a comparison against myself a decade ago.