Hello my name is Goose.
Longtime LessWrong Lurker. Interested in AI Safety, Cognitive Science, Housing Policy, Public Education Policy, and increasingly about communicating issues about AI to non-rationalists/members of the general public.
Hello my name is Goose.
Longtime LessWrong Lurker. Interested in AI Safety, Cognitive Science, Housing Policy, Public Education Policy, and increasingly about communicating issues about AI to non-rationalists/members of the general public.
I also have question2.
IIUC, The method Ryan describes is “generate a bunch of programs” and then “filter those”. He wrote:
> The distribution of programs you are searching over [after generating a bunch of them] has to be pretty close to the right program for Best-of-6k to work at all: if you did best-of-6k for random python programs, this would not work!
I agree that it’s way better than random… but is it better than the human literature? E.g., If you asked me to filter 6,000 NLP papers, I’d expect a great many of them to be “basically true” and a great many more to be “a good effort but ultimately just very wrong”.
Is that the sense in which 6,000 programs he generated and filtered… match the distribution of human output? Like Steven says here, it “spits out tons of confused nonsense with occasional insights, with no labels on which is which. Just like the humans.” Am I thinking about this right?
Thanks for this, as I was reading I thought of this: “I cannot remember the books I have read, anymore than the meals I have eaten; Even so, they have made me”. I used to be very disturbed by not being able to remember all the facts and details I’ve read, and I found this quote quite soothing (I haven’t fact-checked, but it was attributed to Ralph Waldo Emerson).
I interpret questions like “Do you think X will happen?” as meaning something like “Please report your posterior on event X, after all the updates you’ve had so far”. People can answer with just their posterior, or with their posterior and the things that caused updates (or informed the prior, I guess).
If they answer with just their posterior, it might be because the updates aren’t easily verbalizable: e.g., “I spent most of my undergrad learning all the problems with various theories of psychology, and now I have a vague mistrust of social sciences, which I’m only dimly aware of and can’t really defend on any particular point; therefore, I’m a little skeptical of this study but I can’t exactly say why”. Telling me just your posterior is still useful! Because it gives a summary statistic of all the experiences you had, which did update you on this point.
If someone answers with the posterior and their “update events”, this is typically a filtered set of facts. This is because conversations are governed by relevance (speakers tend to cite things that are relevant to the question under discussion), and because the question under discussion is about opinion (in this example), the listed facts are assumed to be part of the opinion (how it came to be). This filtering is not on its own nefarious, because even honest/cooperative speech partners have to filter for relevance to conform to conversational conventions.
Finally, I want to offer an explanation for why this isn’t perceived as rude: listeners still perceive it as an answer to the question about opinion. This is because of the same relevance dynamics mentioned in (2): You’re still answering the question about the opinion. This is similar to why it’s not rude to answer “I’m allergic” when someone asks “Do you like peanuts?”: it’s not technically a statement about preferences, but for all intents and purposes it does answer the question!
This means you can also construct situations where answering with facts would be rude, usually by changing what the conversation is “really” about. E.g. when someone makes a small gesture of ‘social smoothing’ by asking about the weather, and their conversation partner gives a small lecture on meteorological patterns. Sitcoms use this kind of pattern to depict a character as being “nerdy and rude”, for example.
I have two clarification questions.
Question: IIUC the “talk-y part” and “the do-y” part are sub-parts of a single model instance (circuits, or idk some sub-model package of computation). That is, you’re NOT suggesting that the claude app is splitting up work into subagents under the hood, so that 3 models in a trench coat appear to be just 1 model. Do I have that right?
Question: How could this have turned out any different? What else could you reasonably expect to find? Is the alternative “a thinking thing that is aware of all its computations, at all granularities, and can verbalize all of that at any time?”.
If you want to do more than one kind of complicated reasoning (like code AND use English), don’t you need to specialize your parts a bit? I think I’m confused why you’d expect all of those specialized bits to be easy to verbalize, or why you’d expect them to have really direct connections to the language-y bits? Or, is the surprise just how much of the do-y bits don’t directly interface with the talk-y bits?
I agree with basically everything in this comment, except maybe this part: “The emotional impact makes it less easy to treat your belief in it as a tool for impressing your friends and irritating your enemies, rather than something to actually worry about and try to solve if it’s real.”
I just don’t think the argument goes through: “The emotional impact being stronger, makes it LESS easy for you to treat your belief as a tool for impressing friends/irritating enemies”. (Citing HPMOR: “For it is a sad rule that whenever you are most in need of your art as a rationalist, that is when you are most likely to forget it.”)
Also, I think the emotional impact of climate change was, for some people more extreme, and proximal/immediate; definitely more than how they experience AI risk. Especially people who lived in places where the ecosystem was more fragile/environmental damage happened in a rapid and extreme visible way (coral reefs for example changed in ways that were rapid even by human standards?) and who have only just started thinking about “maybe AI will take my job”.
Still, just more reason to stay focused, and fight the polarization as best we can!!