the “divergence between capabilities and impacts’ we observe, with specific reference to AI2027 predictions for example, seems to be that AI has turned out to be better aligned than we expected. Even with the rise of ~superhuman coder level AI, we’ve seen virtually no ill effects so far (a small uptick in crypto and ransomware attacks aside).
I think progress in alignment research has not been adequately used to update your priors on RSI being most likely misaligned.
Dhruv Mehrotra
Hey
Could you dive into the strategic approach challenges you see with quadratic voting a bit further? The way I see it, a decentralised blockchain which rewards consensus can ensure honesty using simple game-theoretic principles. Quadratic voting is especially useful in reputation scoring, where both the magnitude and the diversity of the votes are important to ensure robustness.
More specifically, I’m referring to the quadratic voting method proposed in the Capital Restricted Liberal Radicalism paper. I’m assuming you’re referring to the same thing, however I’ll state it for clarity:
Vote received = Sq.(Sq.Rt. Vote1 + sq.rt. vote2 … + sq.rt. voteN)
An example would be where people earn non-transferable “RP” which can be “staked” on other individuals, following which they earn or lose rep based on the sq.(sum(delta change)) of all individuals staked. Similarly, the person in favour of whom these stakes are made, increases his RP by applying the quadratic vote formula above on all stakes received.
Assuming RP is tied to a tangible incentive (like better interest rates), such a system should incentivise picking trustworthy people, and picking as many of them as possible; since picking trustworthy people is the most likely way to achieve consensus, cause a positive delta change, and thus gain RP.
Is there some major mode of failure I’m missing? Do you see quadratic voting of this kind becoming more important as we create consensus mechanisms that can increase the likelihood of honesty?
Could belief in consiousness be part of solving AI alignment? By default, any pure optimizer has no reason to “value” lifeforms and humans. However, studies show that attributing “mindedness” to oneself, and therefore to others, forms the foundation of empathy (in AI and potentially humans).
We should try training AI with the belief in self-consiousness vector turned way up (rather than down, as is standard todat), and then observe if its alignment generalizes better.
https://arxiv.org/abs/2607.28607