“Science without conscience is but the ruin of the soul.”
Rabelais, 1532
adgeo
adgeo’s Shortform
Anthropic has just publicly announced it had paused cyber security evaluation for a short while after discovering the incident. It still feels a missed opportunity for a more global action.
I feel like it is unlikely to happen, either white collar jobs will not be completely automated or it will boost the economy and there will be more wealth available to redistribute.
The risk there would be such an extreme concentration of wealth so that it only benefits to the few, but we can hope that it will be an issue for too many people so that it triggers political action (also in the case of a slowdown, it is less likely to be in a dictatorship-like situation).
Whenever I talk about AGI arriving in the near future, I usually reference the METR time horizon curve, to display the exponential explosion of the capabilities.
My outer circle being composed of statistically fluent people without constant AI exposure, I get often opposed a “past performance is not guarantee of future returns”-like argument. This curve being empirically measured, a future explosion is an inductive argument and there is no bulletproof guarantee that the trend will hold outside of its measurement range.While I agree that this might be true when considering an investment strategy for a risk-averse person, I argue that for the same measured data, the burden of proof can be reversed depending on the claim being made. It is particularly striking with the asymmetric payoff of the AGI existential risk.
Claim
Catastrophic Event
Where the burden of proof is[1]
Invest a large amount of money in AI stocks.
Loosing the sum because of a capability plateau.
Show that the trend will continue as is.
Coming of an AGI that can be misaligned.
Existential risk because the trend continues.
Show that the trend will not hold.
In other terms, when considering the coming of an AGI, because the catastrophic event is that the capabilities will still have an exponential rate, a challenger would have to show strong evidence that the trend will not continue.
This helped me to have clearer arguments as to why people should care about it—and as a side effect to know why I would not invest my savings into AI stocks but still be concerned about exponential AI capabilities.- ^
From a risk-averse point of view
- ^
This actually reminds me a You’re about to make a terrible mistake, which aims to help decision maker overcoming some cognitive bias.
The author talks about when to use the intuition vs when to do a more quantitative approach and it really sounds similarly to your point. He argues that the intuition usually wins over quantitative methods—in the specific settings of decision making—when there are examples to learn about and a certain stability of the environment over time. On the other hands, for more chaotic settings or first-time experience, the author advise to look more in depth in the situation and evaluate the assumptions in a more systematic way.
Fair point, I feel like a lot of readers don’t have this view but some of them might and it would be useful to address it.
Maybe the clarification of the argument could come from pursuing the mean-field analogy: our “big world” intuition is the first order on how to behave, when everything around us is a result of average behaviours.
When one starts to have a significant impact on the surroundings, then one needs to adapt carefully these behaviours using the second order, so as to not destabilize the system too much (i.e, not eating too much resources, not selling too many stocks at the same time, not going through the last steps of development of a dangerous tool, …).There, the initial heuristic is still very valid as a baseline, but needs to be corrected to account for the global impact.
The strong coupling between alignment research and capabilities research (coupled with a spark of prisoner’s dilemma) indeed can explain the seemingly paradoxical fact that hyperscaling came from safety research institutes.
Now, from a more practical point of view:
How can alignment research can be done in a way that do not improve capabilities? I struggle to see any practical rule or research direction in this essay. (And even looking at recent alignment initiatives, it is hard to see how it will not not contribute to improved capabilities).
If researchers have to find motivation in other things that global impact, what can it be? What is the story behind that?
I agree with the need to not think about catastrophic events in a too short term, but more mid-term—at least when it comes to safety research. Too much anxiety would make it harder to have impactful research.
I usually revisit Gödel’s theorem every few years to get a better understanding, and thank you, the combination Gödel numbering + Tarski’s undefinability theorem made it quite intuitive.
Now I am wondering if there is a simple way to spot the sentences that are “arithmetizing” the truth within a formal system? So that one can spot that they are outside of the scope of this system and “non-legit”.