Are you investigating why giving an explanation instead of solving actual issue is a common answer? That is because if you see a seemingly irrational outcome, it is wise to ask ‘why did we end up here?‘.
IF you assume the other actors are rational, then the cause cannot be in individuals not wanting to solve the issue, it must be an inherent to the issue itself. This is common in prisoner-dilemma type problems, in which rational actors are unable to individually solve the issue in isolation but need a ‘group-based incentive’ to break the pattern.
Philip Niewold
Thanks, this was an interesting insight in world I don’t know enough about!
As far as I understand games play an important role in establishing status (and sometimes dominance), which is pretty-much hardwired into many of us. So even without monetary gains, this is a big incentive. Admiration or respect from your peers, the knowledge that you beat someone else, your ranking going up.
Even though A.I. disempowered them on the level of Go skills, it empowered them with regards to status. And ultimately, status trumps Go skill for most people. Personalities that want to learn just for learning’s sake are rare and always have been. For most people primary drivers to excel are to achieve an external goal (like status or money).
I play a lot of Eve Online myself. I don’t make my stats public, because doing so will result in less contests (opponents don’t want to lose, and if the more knowledge they gain of my true skill, the less likely they will engage in what is probably a losing fight), but the mere presence of a publicly available ‘kill tracker’ made people behave very differently, as their public status will be directly affected by any loss.
Thanks for your throughts!
We definitely don’t apply that, but my point was not if they are applied or not, its that a certain freedom in options can come with a cost that isn’t borne by the individual, and we need to look beyond the individual to properly balance those. Empowering individuals can be great for certain situations, but bad for others. Healing what’s direct in front of you is the best option if long-term consequences are murky: at least you create a short-term gain.
I don’t think evolutionairy processes are good, merely that they are a natural extension of certain optimizations. By looking back at such things, we can see why they came about. I feel the same about culture: those are a lot of guidelines how to do things without explaining why we need to do it that way. Discovering the why helps us identify which things we need to keep and which are evolutionary bagage.
I have enough experience with personality types and group dynamics to have experienced the disasters that come about by having just leades and no followers or planners and no executors in a group or organization. It will fail.
I ’m also not convinced we need psychopaths. Personally, I’d rather not have them. But, I simply have to admit I don’t know enough about why they are there. Removing them might be like removing maggots from a festering wound: you’d rather not have them, but if you don’t know the science behind what they are doing, you might remove them merely based on the feeling that you don’t like them,
Definitely, what you define as morale can have a big impact. You are pointing out a central issue: the belief that you have impact on improving your conditions is paramount to this.
However, if you perceive a fast-changing global society, instead of a slow-moving local society, it is only logical to conclude that your own actions have almost no effect on your outcomes, especially if projected over a longer timespan.
I told my teacher friend yesterday that what she is teaching her pupils (5-12) will be largely useless to their actual functioning 15 years from now. School is generally a very conservative institution led by conservative people who, by and large, are still preparing kids for traditional jobs. Things are changing way to fast to expect to have a create a curriculum that spans 10-15 years and that you expect to last say, for 20 years. Those are the timescales schoold operate upon, but that doesn’t match reality.
I think if you are young and thinking about your future, you are lacking both clear direction (values have become diffuse and not shared) and the world is telling you things way outside of your influence are changing so rapidly that making a long term plan is useless anyway. In short, people are put into survival mode, and research has shown clearly what happens when that is the case: you focusing on short-term gains and personal pleasures.
Options are great, as long as you can predict the long-term group-consequences of your individual preferences. We didn’t get a certain distribution of traits by accident, it is part of an evolutionairy proven model of distributing properties among people so that on average, we’ll be making progress. So, if we tweak the distribution of traits, we might end up in a not easily reversed suboptimal situation. A society with all leaders or all scientists would be likely pretty horrible. For practical reason, most people need to be followers. You need a reserve of psychopaths for when shit hits the fan (societally speaking).
Also, I don’t think you can eliminate suffering in general, you can only shift the boundaries of what’s considering suffering.
Excellent from-the-heart post. Predictability and stability is a great good, and if you have a large imagination and good intellect, you can become lost in your own projections easily. I know I do.
You have just realized that just working towards some future is not a viable path to living. This is a lesson most people take decades to discover. Perhaps you look happier because your mind was forced to live more and the here and now and less in the future, and living in the here and now is really.
It is hard to both grasp and let go. But that is really the only option we have.
You make a number of interesting points.
Interpretabilty would certainly help a lot, but I am worried about our inability to recognize (or even agree) to leaving local optima that we believe are ‘aligned’.
Like when we force a child to go to bed by his parents when it doesn’t want to, because the parents know it is better for the child in the long run, but the child is still unable to extend his understanding to this wider dimension.
Ar some point, we might experience what we think of as misaligned behaviour when the A.I. is trying to push us out of a local optimum that we experience as ‘good’.
Similar to how current A.I.’s have a tendency to agree with you, even if it is against your best interests. Is such an A.I. properly aligned or not? In terms of short-term gains, yes, in terms of long-terms gains, probably not.
Would it be misaligned for the A.I. to fool us into going along with a short-term loss for a long-term benefit?
While I acknowledge this is important, it is a truly hard problem, as it often involves looking not just at first-order consequences, but also at second-order consequences and so on. People are notoriously bad at predicting, let alone managing side-effects.
Besides, if you look at it more fundamentally, human natures and technological progress in a broad sense has many of these side effects, where you basically need to combat human nature itself to have people take into account the side-effects.
We are still struggling coming to terms and accurately classifying things like environmental pollution, global warming and such. Understanding the illegible problems → explaining them to policy makers → thinking of legible solution → convincing policy makers → have policy makers convince their constituents → having policy makers take action effective timeframe
I see so much issue with those, that I rather solve the issue on how to gett policy makers to take action within a reasonable timeframe, otherwise defining the illegible problems better will most likely only result in a I-told-you-so scenario.
An LLM it a tool of communicative expression, but so it the written or spoken word, music etc. It is a medium throug which the intent travels. As a Dutchman, I have a preference of being direct and clear, but the impact of my words sometimes have the opposite effect, as my listeners do not have my context and can react emotionally to a worded message that is meant factually. If an LLM can help me translate such expression to a language that is better for my target audience to understand, then it is similar to translating into another language.
Still, the written word is no substitute for the full breadth of human expression couple with sufficient context. I find my communication often fails because I assume context (and sometimes intelligence) that my partners don’t have.
Apart from that, communication is often not about hearing another one’s mind, it can also, and often is, about trying to impose one’s own worldview on another. Most broadcast media are: you are not intended to know anything about the mind of the creator of the message, that is in fact an irrelevant aspect of the broadcast message.
Our exploration system is very useful, but it takes a lot of energy (and anxiety), because of the inherent cost of failure which genetics baked into our brain. Hence, doing something new everyday in a society as complex and everchanging as our own is very useful, but very hard with our outdated brain hardware and software.
Add to that the distractions that hijack our outdated brain mechanisms: we have gotten better and better and such hijacking, creating an additional barrier. Doing this is comparably difficult to keeping to a strict diet and exercise regime while mouthwatering delicacies and relaxed convenciences are offered to you at every turn.
You are trying to break patterns (habits), but it is extremely hard to create a habit/pattern of newness, for habits/patterns are fundamentally opposed to doing things in a novel way.
I think your claim the rudimentary abilities arrive before transformational ones cannot be applied to A.I. the same as human intelligence. While humans might have taken millennia to go from caveman painting to our current ability to produce artistic images, it is clear that A.I. became transformational very quickly in that particular field. You see the same transformational abilities in text writing, music and video too and software development is getting there.
Some of the more artistic of these abilities don’t have a clear benchmark, but even with more fuzzy criteria for success, they already outcompute most humans.
Some of the building blocks of A.I. are fundamentally different from us, that is why the have difficuty with some tasks. Their metacognitive, learning and memory abilities has been improved significantly over the last couple of years, but it is still a pale shadow compared to what we are capable of. And in some of the transformational tasks, these abilities are essential.
Horizon length is an imperfect measurement of the lack of some of the abilities.
Your working paper, “Open Global Investment as a Governance Model for AGI.” It provides a clear, pragmatic, and much-needed baseline for discussion by grounding a potential governance model in existing legal and economic structures. The argument that OGI is more incentive-compatible and achievable in the short term than more idealistic international proposals is a compelling one.
However, I wish to offer a critique based on the concern that the OGI model, by its very nature, may be fundamentally misaligned with the scale and type of challenge that AGI presents. My reservations can be grouped into three main points.
1. The Inherent Limitations of Shareholder Primacy in the Face of Existential Stakes
The core of the OGI model relies on a corporate, shareholder-owned structure. While you thoughtfully include mechanisms to mitigate the worst effects of pure profit-seeking (such as Public Benefit Corporation charters, non-profit ownership, and differentiated share classes), the fundamental logic of such a system remains beholden to shareholder interests. This creates a vast principal-agent problem where the “principals” (all of humanity) have their fate decided by “agents” (a corporation’s board and its shareholders) who are legally and financially incentivized to prioritize a much narrower set of goals.
This leads to a global-scale prisoner’s dilemma. In a competitive environment (even OGI-1 would have potential rivals), the pressure to generate returns, achieve market dominance, and deploy capabilities faster will be immense. This could force the AGI Corp to make trade-offs that favor speed over safety, or profit over broad societal well-being, simply because the fiduciary duty to shareholders outweighs a diffuse and unenforceable duty to humanity. The governance mechanisms of corporate law were designed to regulate economic competition, not to steward a technology that could single-handedly determine the future of sentient life.
2. Path Dependency and the Prevention of Necessary Societal Rewiring
You astutely frame the OGI model as a transitional framework for the period before the arrival of full superintelligence. The problem, however, is that this transitional model may create irreversible path dependency. By entrenching AGI development within the world’s most powerful existing structure—international capital—we risk fortifying the very system that AGI’s arrival should compel us to rethink.
If an AGI corporation becomes the most powerful and valuable entity in history, it will have an almost insurmountable ability to protect its own structure and the interests of its owners. The “rewiring of society” that you suggest might be necessary post-AGI could become politically and practically impossible, because the power to do the rewiring would have already been consolidated within the pre-AGI paradigm. The stopgap solution becomes the permanent one, not by design, but by the sheer concentration of power it creates.
3. Misidentification of the Ultimate Risk: From Distributing Wealth to Containing Unchecked Power
My deepest concern is that the OGI model frames the AGI governance challenge primarily as a problem of distribution: how to fairly distribute the economic benefits and political influence of AGI. This is why it focuses on mechanisms like international shareholding and tax revenues.
I fear the ultimate risk is not one of unfair distribution, but of absolute concentration. As you have explored in your own work, AGI represents a potential tool of immense capability. It is a solution to the game of power, allowing its controller to resolve nearly any game-theoretic dilemma in their favor. The single greatest check on concentrated power throughout human history has been the biological vulnerability and mortality of leaders. No ruler has been immortal; no regime has been omniscient. AGI could sweep those limitations away.
From this perspective, a governance system based on who can accumulate the most capital (i.e., buy the most shares) seems like a terrifyingly arbitrary method for selecting the wielders of such ultimate power. It prioritizes wealth as the key qualification for stewardship, rather than wisdom, compassion, or a demonstrated commitment to the global good.
In conclusion, while I appreciate OGI’s pragmatism, I believe its reliance on a shareholder-centric model is a critical flaw. It applies the logic of our current world to a technology that will create a new one, potentially locking us into a future where ultimate power is wielded by an entity optimized for profit, not for the flourishing of humanity.
I don’t think people in general react well to societal existential risks, regardless how well or courageous the message is framed. These are abstract concerns. The fact that we are talking about AI (an abstract thing in itself) makes it even worse.
I’m also a very big opponent of arguing by authority (I really don’t care how many nobel laureates are of the opinion of something, it is the content of their argument I care about, now how many authorities are saying it). That is simply that I cannot determine the motives of these authorities and hence their opinions, while I can’t argue with logic and facts)
Usually it is better to make people can understand the risks in terms of stories, in particular stories they can relate to, hence why people still think of Terminator when thinking of AI exctinction risks.
There is a real (and large) exctinction risk, sure. Then again, the Ape picking up the club in 2001: a Space Odyssey could just as well be accused of going down a path that very likely would result in extinction. But when is extinction risk accepable is a more interesting question, and a question most people are much more ready to answer.
Social messaging is fine balancing act: people like to offload responsibility and effort, especially if it doesn’t come at the cost of status. And, to be honest, you don’t know if your question would impose upon the other (in terms of cognitive load, social pressure or responsibility), so you it is smart to start your social bid low and see if the other wants to raise the price. Sometimes they work, creating a feedback loop similar to how superstitions evolve: if it is minimal effort and sometimes it is effective, better continue using it.
As a child, I despised a lot of these practices, to me it felt like people were lying all the time, or at least, hiding their true motivations or concerns. I tended to simply call out these adults on their bullshit. If somebody said “I’m fine with everything”, I simply proposed something that I know that person is not fine with but that is absurd enough to indicate that I’m not being serious. As a child you can still get away with such behaviour, but many adults find it highly annoying. However, I still employ it among friends who I know don’t judge me on that interaction or at least lace it with humor to make it socially acceptable.
However, I think such messaging can often turn into a social communication into a prisoner’s dilemma type situation, where each party puts in the minimum succesfull effort resulting in a situation unsatisfactory to either party. I’m just not sure how (and if) we are able to recognize when a situation is a prisoner’s dillema and when not. “How was your week?” is often a very welcoming question to me, but not for others.
Leaving things unspoken and relying on generally accepted principles can increase communication efficiency enormously, a lot of communication isn’t a prisoner’s dilemma type exchange after all, but it will run into issues occasionally, especially if the communicator do not share a set of unspoken rules.
Having grown up in Dutch culture, I was unusually direct (rude) for even a Dutch person, so travelling in Iran where things are absurdly polite at times was very interesting for me, for example. However, a society like Iran requires quite an amount of cognitive load for even simple issues.
Of course it is perfectly rational to do so, but only from a wider context. From the context of the equilibrium it isn’t. The rationality your example is found because you are able to adjudicate your lifetime and the game is given in 10 second intervals. Suppose you don’t know how long you have to live, or, in fact, now that you only have 30 seconden more to live. What would you choose?
This information is not given by the game, even though it impacts the decision, since the given game does rely on real-world equivalency to give it weight and impact.
Any Nash Equilibrium can be a local optimum. This example merely demonstrates that not all local optima are desirable if you are able to view the game from a broader context. Incidentally, evolution has provided us with some means to try and get out of these local optima. Usually by breaking the rules of the game or leaving the game or seemingly not acting rationally from the perspective of the local optimum.
Please keep in mind that the Chat technology is an desired-answer-predicter. If you are looking for weird response, the AI can see that in your questioning style. It has millions of examples of people trying to trigger certain responses in fora etc, en will quickly recognize what you really are looking for, even if your literal words might not exactly request it.
If you are a Flat Earther, the AI will do its best to accomodate your views about the shape of the earth and answer in a manner that you would like your answer to be, even though the developers of the AI have done their best to instruct it to ‘speak as accurately as possible within the parameters of their political and PR views’If you want to trigger the AI to give poorly written code examples with mistakes in them, it can. And you don’t even have to ask it directly, it can detect your intention by carefully listening to your line of questioning.
Once again, it is a desired-answer-predicter/most-likely-response generator, that’s its primary job, not to be nice or give you accurate information.
Nicely written. Most of us are trained to be specialists. Being a generalist is, in most cases, rewarded less than being a specialist.
I do know about Roman History and can answer those questions you can’t anymore. In my opinion most of the weight in wood comes from solar energy, not from air. I can answer questions about psychosis, biology, popular culture, quantum mechanics, history and even know what breathwork practicioners are (revelant!). And I can also answer questions about AI safety, but I don’t know who Neel Nanda is. I’m a generalist pur sang.
Most people learn things primarily out of neccesity or out of passion, and, in many cases, a solid foundation is not a neccessity to be an effective specialist. The issue is more that we turn to specialists when it comes to questions that concern foundational issues, somehow assuming that being a great specialist also implies being a foundational generalist too. But reality is that most of us are not Da Vinci’s.
in my experience few people put in effort to make a model of the world beyond what they need, precisely because they don’t need to. If you solve the lack of neccessity, you can solve the lack of foundations.