I’m a computational cognitive scientist studying decision-making and introspection in humans and LLMs. I earned my PhD at Harvard in 2022, and have been a postdoc at Princeton since then.
Adam Morris
This is haunting.
I’ll say that I genuinely found this funny (although I’ve been told that I’m easy to please humor-wise 🙂)
EDIT: Okay now that I’ve finished it, I also found it moving! Again, I’m probably easy to please. It’s clearly not as good as the rest of your fiction, but it’s genuinely a better story than I could write.
Fascinating point, I think you’re right. Just to repeat your point in my own words: The problem is that, if the activation steering makes the model want to talk about the injected concept, and if it knows that saying “yes, I received an injection” will give it a chance to talk about the concept later in the response, then it will say “yes” in order to talk about the concept later (even if it actually had no metacognitive awareness of the injection). Is that what you’re saying?
Tests of LLM introspection need to rule out causal bypassing
Is there reason to think that Bores or Wiener are not trustworthy or lack integrity? Genuine question, asking because it could affect my donation choices. (I couldn’t tell from your post if there were, e.g., rumors floating around about them, or if you were just using this as an example of a key question that you thought was missed in Neyman’s analysis.)
Self-interpretability: LLMs can describe complex internal processes that drive their decisions
Got it. Okay thanks!
Earnest question: For both this & donating to Alex Bores, does it matter whether someone donates sooner rather than a couple months from now? For practical reasons, it will be easier for me to donate in 2026--but if it will have a substantially bigger impact now, then I want to do it sooner.
One small suggestion: When I read this, I genuinely couldn’t tell whether “Gray swans: None detected this week” was a joke (like you were pretending to look for literal gray/black swans), or if it meant something serious. After reading your website, my guess is that it’s meant to be serious—but I’m still not sure, and if it is serious then I don’t know what it means. (My understanding is that “black swan” means an unexpected, highly improbable / out of distribution event, so it wasn’t clear to me what it would mean in this context to be generally looking for global gray/black swans.) Might be worth clarifying or finding other terminology, if you want readers like me to quickly grok what you mean.
We haven’t had one yet! But we only did it ~3 times. Obviously people are more careful than they’d normally be while dancing on the slippery floor.
I’ll add to this list: If you have a kitchen with a tile floor, have everyone take their shoes off, pour soap and water on the floor, and turn it into a slippery sliding dance party. It’s so fun. (My friends and I used to call it “soap kitchen” and it was the highlight of our house parties.)
Printable book of some rationalist creative writing (from Scott A. & Eliezer)
I see, that makes sense. Thank you!
Can you help me see this point? Why not correct it in the dataset? (Assuming that the dataset hasn’t yet been used to train any models)
I’m long overdue here, but thank you so much for doing this!! I’ve been wanting this for a long time and just discovered this post :)
see my comment above—I (ironically) meant aphasia
hahaha I actually also meant aphasia :P
This is ~even more~ anecdotal, but me and several of my friends have noticed increased anosmia since the pandemic, but critically starting before any of us got covid (and including friends who never got it). We conjectured that it could be from some combination of very high stress levels for a long time + social isolation? Just to add some data points to the mix.
I know there’s this kind of evidence, but it’s so discordant with my experience. I use Claude Code/Cowork all day every day for work, and haven’t once it had do anything reward-hackey or strategically deceptive (at least that I can remember? or that I’ve caught?) in the last ~3 months. Do other people actually have the experience of it reward hacking in their day-to-day use?