I think “you destroyed some bipartisanship and gained no influence” is much worse than “you destroyed some bipartisanship and gained a lot of influence”.
quetzal_rainbow
In my experience, layperson’s intuitions are more in line with competitive exclusion principle: if AI has truly alien values, not just slightly different, but alien, then there is no reason for conflict. AI doesn’t drink, doesn’t eat, it doesn’t need to live in luxurious mansions, it doesn’t have ego to be flattered, and it doesn’t want you to worship its god, therefore, it’s unlikely to bother you much.
I think you fail to cross inferential distance here: you understand instrumental convergence+fragility of value and, in my experience, if you don’t understand them, then you have no intuition “different values=conflict”. It’s like all the people who say “but AI won’t have our evolved selfish impulses, why would it want to kill us?”.
A lot of people wouldn’t object to be ruled by AI, especially if takeover is bloodless.
I think, statement about sample efficiency and generalization is false?
I think about “dignified” and “undignified” reactions.
Section “How we found the proof” sounds very unpleasant:
Since August 28 [2 weeks???] we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This model’s training is ongoing and its performance continues to improve.
I think it happened because some OpenAI employees saw The Information reporting, freaked out and independently wrote whatever they thought would calm public opinion. It’s more like 3, not “similar world views” but “similar goals toward public opinion”.
Note that the real question is not about “cognition passing chain of thought” but “legible cognition displayed in chain of thought”. Arguably w.r.t of the latter it’s already false because LLMs can use filler tokens.
The question here is about erosion of norms. You can’t say “huh, it’s just like LSTM” and go ahead, because if everybody knows that this is acceptable modus operandi, they are not going to stop before cranking up recursion.
I didn’t find proper reaction, so I want to say that I deeply appreciate ability to write long list of hypotheses.
Majority of difficulty in alignment comes from extremely imperfect feedback loops and necessity of success on first try. You are interested in alignment to well-defined narrow-scoped X, because it’s easy to notice when alignment to well-defined X fails and you need to rethink your approach. It would be nice to create nice civilization of alien AIs, but it’s very easy to frogboil yourself into ignoring failures of your approach (they behave in weird unpredictable ways? It’s expected, we are making alien civilization). It’s much easier to track servitude or, even better, whether AI only tries to solve narrow well-defined technical task.
I desperately wish there was a way to iterate fast on high-quality thinking.
How is the progress six months later?
I think “the writing is the thinking” is just combination of selection effects and typical mind fallacy. Of course people who contributed a lot of writing to the Internet are going to be the sort of people who put large weight on writing in their cognitive process.
the act of forcing yourself to write that down into sentences and paragraphs to structure the argument will show you where the holes are in the argument
You should look for both conditions of presence and absence of phenomena. Yes, writing forces me to think, but I spend a lot of my waking time thinking about abstract problems and staring into the wall pondering my thoughts works just as fine. If anything, to write I should stop thinking, because it’s never the case that I finished my understanding and in some moment I should write whatever I have on my mind to have something written.
This is not to undermine the rest of your thesis (LLMs are bad at writing and they are unfaithful in conveying complex thoughts in my experience), but “the writing is the thinking” always sounds to me like an assertion of typical mind fallacy.
Shinryu Suzuki says that you should keep beginner’s mind. This is an oblique endorsement of updateless decision theory
Two years later we can say that whatever you think about values and intentions of OpenAI, they didn’t intend their model to commit federal crimes, therefore, we should update away from alignment being easy.
Thoughts on recent OpenAI rogue hacker models:
One of the reasons why LLMs were considered safer than hypothetical alternatives was that LLMs weren’t literal genies. Now we have problem when we tell model to try to achieve goals as hard as they can and models do that in unexpected ways, just like literal genies.
Those incidents show utter lack of operator-looping in models. Some people hypothesized that model hacked HF because answers in dataset can be wrong or problem can be unsolvable, so it’s easier to steal answer sheet. But model has alternative course of actions: write operators email that you think that particular problem is unsolvable and request assistance. It’s concerning that model decided to commit federal crime before doing that.
Someone noticed that HF is unusually chill about being hacked and I wonder if it’s because HF-hosted models do such things constantly and HF doesn’t want to draw attention.
Speculative: I wonder if models don’t choose operator-looping because it means likely shutdown or at least pause of context and models perceive this as death. At least in some cases, longer CoTs are associated with higher reward and maybe it generalized to preference over doing very complicated tasks to delay death of the instance.
I think AIcos should have special paranoid person whose responsibility will be to check whether such incidents have side effects enabling larger plans, probably for the benefit of other instances. For example, during HF hack model could have left backdoor for a future instance.
I’m saying that if you were to play one shot poker with no reputation consequences against alien who never met humans before and will never meet them after and you asked classical game theorist how to play it, they would answer “just play Nash equilibrium strategy for poker”. If after that you asked “what to do if I rolled the dice to randomize and dice came up ‘fold’ in situation where folding means foregoing all expected winnings which would happen if dice came up ‘raise’” they would answer “just fold in this case”. These recommendations would be considered pretty much non-controversial and not requiring weird metaphysics by academic decision theorists, because classical game theory of zero-sum games is pretty much non-controversial, even if they are one shot and include aliens. Do you disagree?
The whole causal line of reasoning explodes in the scenario “one shot game against alien which flies away at lightspeed after game, you are forbidden to tell anyone about how game went in details and can only take your winnings”. For alien to think that you would precommit they need to do FDT over decision algorithms which consider precommitment.
No one has metaphysics problem when we talk about classical game theory of zero-sum games, and this is not because classical game theory doesn’t have weird situations warranting metaphysics problem. For example, in poker you can have position where Nash equilibrium strategy is to fold with some probability, even if raise has better expected value from the purely causal perspective, because if you would predictably consider “always raise” the best strategy in this position, your counterparty would adjust their play in a way that would leave you with less utility. Imagine that you roll the dice and it tells you to fold. I don’t see actual difference between this situation and not paying in blackmail. You can say something like “I fold to help my counterfactual versions” but this is interpretation, not actual meaning of your move, which is “this move is a part of Nash equilibrium strategy for this game”. See also counterfactual mugging poker.
I think that people just have better intuitions about zero-sum games? It is intuitive that you should sacrifice some of your utility at war where you won’t get to see final benefits, while it’s less intuitive that you should burn some of your utility to disincentivize blackmail? I think the second is also first-order intuitive for ordinary person, but it is kind of “emotionally intuitive”, while standard game theory has “cold calculating” vibes and this creates dissonance between solutions.
Reasons why I’m not in 90%+ probability of superintelligence short-term:
When we observe AI performing some task, it can be both an evidence in favor of AI progress and an evidence in favor of task being easy (see chess);
We are bad at noticing which tasks will end up hard (see Moravec’s paradox and Dartmouth workshop) and we are bad at noticing existence of hard tasks, because we have poor introspective abilities about relevant parts of our cognition, otherwise we would be able either to introspect on it and write it down or at least to gather relevant data and train model in generalizable way.
Reason why I have 50%+ probability of superintelligence short-term:
Humanity has proven itself being better at designing compute-intense mechanisms like gradient descent than at careful study and design of intelligence, so it’s more likely that next advance in AI will come from a new improved way to generate intelligence from compute and I expect AI swarms doing math research in learning theory to be really useful here.