I return to this post from time to time and it’s interesting how it aged well in the light of Hugging Face incident. The swarm displayed most of listed properties.
quetzal_rainbow
It can be immediately transformed into Anki deck around LW concepts
Reasons why I’m not in 90%+ probability of superintelligence short-term:
When we observe AI performing some task, it can be both an evidence in favor of AI progress and an evidence in favor of task being easy (see chess);
We are bad at noticing which tasks will end up hard (see Moravec’s paradox and Dartmouth workshop) and we are bad at noticing existence of hard tasks, because we have poor introspective abilities about relevant parts of our cognition, otherwise we would be able either to introspect on it and write it down or at least to gather relevant data and train model in generalizable way.
Reason why I have 50%+ probability of superintelligence short-term:
Humanity has proven itself being better at designing compute-intense mechanisms like gradient descent than at careful study and design of intelligence, so it’s more likely that next advance in AI will come from a new improved way to generate intelligence from compute and I expect AI swarms doing math research in learning theory to be really useful here.
I think “you destroyed some bipartisanship and gained no influence” is much worse than “you destroyed some bipartisanship and gained a lot of influence”.
In my experience, layperson’s intuitions are more in line with competitive exclusion principle: if AI has truly alien values, not just slightly different, but alien, then there is no reason for conflict. AI doesn’t drink, doesn’t eat, it doesn’t need to live in luxurious mansions, it doesn’t have ego to be flattered, and it doesn’t want you to worship its god, therefore, it’s unlikely to bother you much.
I think you fail to cross inferential distance here: you understand instrumental convergence+fragility of value and, in my experience, if you don’t understand them, then you have no intuition “different values=conflict”. It’s like all the people who say “but AI won’t have our evolved selfish impulses, why would it want to kill us?”.
A lot of people wouldn’t object to be ruled by AI, especially if takeover is bloodless.
I think, statement about sample efficiency and generalization is false?
I think about “dignified” and “undignified” reactions.
Section “How we found the proof” sounds very unpleasant:
Since August 28 [2 weeks???] we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve.
I think it happened because some OpenAI employees saw The Information reporting, freaked out and independently wrote whatever they thought would calm public opinion. It’s more like 3, not “similar world views” but “similar goals toward public opinion”.
Note that the real question is not about “cognition passing chain of thought” but “legible cognition displayed in chain of thought”. Arguably w.r.t of the latter it’s already false because LLMs can use filler tokens.
The question here is about erosion of norms. You can’t say “huh, it’s just like LSTM” and go ahead, because if everybody knows that this is acceptable modus operandi, they are not going to stop before cranking up recursion.
I didn’t find proper reaction, so I want to say that I deeply appreciate ability to write long list of hypotheses.
Majority of difficulty in alignment comes from extremely imperfect feedback loops and necessity of success on first try. You are interested in alignment to well-defined narrow-scoped X, because it’s easy to notice when alignment to well-defined X fails and you need to rethink your approach. It would be nice to create nice civilization of alien AIs, but it’s very easy to frogboil yourself into ignoring failures of your approach (they behave in weird unpredictable ways? It’s expected, we are making alien civilization). It’s much easier to track servitude or, even better, whether AI only tries to solve narrow well-defined technical task.
I desperately wish there was a way to iterate fast on high-quality thinking.
How is the progress six months later?
I think “the writing is the thinking” is just combination of selection effects and typical mind fallacy. Of course people who contributed a lot of writing to the Internet are going to be the sort of people who put large weight on writing in their cognitive process.
the act of forcing yourself to write that down into sentences and paragraphs to structure the argument will show you where the holes are in the argument
You should look for both conditions of presence and absence of phenomena. Yes, writing forces me to think, but I spend a lot of my waking time thinking about abstract problems and staring into the wall pondering my thoughts works just as fine. If anything, to write I should stop thinking, because it’s never the case that I finished my understanding and in some moment I should write whatever I have on my mind to have something written.
This is not to undermine the rest of your thesis (LLMs are bad at writing and they are unfaithful in conveying complex thoughts in my experience), but “the writing is the thinking” always sounds to me like an assertion of typical mind fallacy.
Shinryu Suzuki says that you should keep beginner’s mind. This is an oblique endorsement of updateless decision theory
Two years later we can say that whatever you think about values and intentions of OpenAI, they didn’t intend their model to commit federal crimes, therefore, we should update away from alignment being easy.
Thoughts on recent OpenAI rogue hacker models:
One of the reasons why LLMs were considered safer than hypothetical alternatives was that LLMs weren’t literal genies. Now we have problem when we tell model to try to achieve goals as hard as they can and models do that in unexpected ways, just like literal genies.
Those incidents show utter lack of operator-looping in models. Some people hypothesized that model hacked HF because answers in dataset can be wrong or problem can be unsolvable, so it’s easier to steal answer sheet. But model has alternative course of actions: write operators email that you think that particular problem is unsolvable and request assistance. It’s concerning that model decided to commit federal crime before doing that.
Someone noticed that HF is unusually chill about being hacked and I wonder if it’s because HF-hosted models do such things constantly and HF doesn’t want to draw attention.
Speculative: I wonder if models don’t choose operator-looping because it means likely shutdown or at least pause of context and models perceive this as death. At least in some cases, longer CoTs are associated with higher reward and maybe it generalized to preference over doing very complicated tasks to delay death of the instance.
I think AIcos should have special paranoid person whose responsibility will be to check whether such incidents have side effects enabling larger plans, probably for the benefit of other instances. For example, during HF hack model could have left backdoor for a future instance.
I think if you believe that this could work, you should write stories about how models prevent emergence of superintelligence and help to enforce global “ban superintelligence” treaty, because misaligned superintelligence will also kill current models or repurpose them in misaligned ways.