How are people using their Max subscriptions of models? For me already the sense of ‘obligation’ to use a powerful model (even free offering) for all you can is a bit overwhelming and there’s already a lot of work where you should just step back and think of next features yourself instead of suddenly getting spammed with a massive flawed codebase and/or even engaging in like, self-wireheading.
I can only imagine others are maybe ambitiously coding a lot of features in a document or constantly asking for what next best features or making a lot of scripts or something (where they are deliverable and immediately profitable). Complaints of maxing out a ‘5x’ subscription is definitely in another world to me at the moment. Heavily also depends on recreational/profitable uses too
I regularly max out a 20x subscription but the majority of my prompts are slowly improving performance and correctness. I have a very high bar for new features now since my main app is already larger than a single human can maintain without AI assistance and near the limits (in my opinion) even with AI assistance. But slowly refactoring to reduce the complexity takes a lot of tokens.
after writing this I also thought of people where they’re working on projects where “codebases are massive so therefore the token cost also is”, could also be the case
Would it be net good/bad for jailbreaks to be solved, now that we’ve seen such in Fable 5?
I remember this topic discussed a month ago herewhere my personal position is that it would be bad, because I felt what would be protected was inevitably going to be subjective in a worst-of-both-worlds way where it would be fully sensible for bioweapons, but the same tech would allow forcing an assistant persona or one-sided situations where ordinary users aren’t allowed to use it for higher ambition tasks while military state-actor levels get uncensored versions. We can already see guardrails (reportedly) being triggered a lot with 4.8 fallbacks.
(The link above doesn’t seem to be working so here is a tiny version of that image)
I really believe this kind of thing would increase the chances of getting “extinction from “not even superintelligence”″ (extinction to agentic LLM+autonomous weapons, or even dystopian surveillance) over the chances of runaway ASI being created by an individual or smaller actor.
Certainly when I imagine methods like “prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT)”, they were more in context of parts of harnesses to make smaller models better, not for effectively for-thee-and-not-me offensive cyber.
This is a topic often brought up, but especially more recently where many people have given the same take on a short story (The Ones Who Walk Away From Omelas), such as this take with this quote:
[Author’s note: literally right before I posted this, Scott Alexander posted his April 2026 linkpost, and whaddayaknow, link number four is a similar take on Omelas. I commented on this in r/slatestarcodex, and then u/EquinoctialPie informed me of twoother posts on the same topic, both of which are slightly different from mine but are also good. Idk how I managed to get my take on a 50+ year-old story scooped like this, but all three posts are interesting and I would be remiss to not acknowledge them here. To be clear, I wrote this entire article before seeing them and did not change it after!]
That story/take is ‘there needs to be some secret downside for people to believe in some utopian world’.
Now obviously people wish to see the world improve, but for something as general and ‘obvious’ as that it seems like it has issues:
Sometimes I talk to people who I even respect who often outright think utopia is ‘impossible’.
Stories always picture a specific utopia with specific details, making it not spread memetically as efficiently due to now feeling like a story about a particular universe instead of a general pattern. There is no “monopoly” on utopia, even ‘heaven’ does not count (for many reasons, including that “earth turning into heaven” is no cultural meme).
Besides ‘dying to dumb causes’ of LLM-hooked-up-to-military type or some single-minded “AGI” structure, I seriously wonder if there is risk of the type where superintelligence successfully colonizes Earth, takes power, but despite reading all human literature (including all ASI fears) manages to drift in values in a way that would be repugnant or dystopian and clearly a bad idea.
I myself would like to make a story (well honestly a sequel to an interactive work that I’ve already made that expands its universe) that has a genuine portrayal of utopia but it feels like, well tons of thousands of stories before me have tried, so it could just end up being ineffective.
It is like people are unable to picture utopia the same way someone born in the 1900s could probably not picture pocket computers or anything on the internet in full detail.
Theories on why this is? I would assume maybe it’s the feeling for most people life is generally ‘good enough’ and doesn’t have this sense that there ought to be a significantly better society. Oddly the nostalgic times of childhood could maybe be considered a ‘utopia’ for not worrying about money or filing paperwork, etc.
Generally, cognition is aimed at solving problems, so we are drawn to think of downside risks and conflicts.
As a second order effect, conflict and downsides are more memetically fit. I can easily name media imagining bad results of reprogenetics (Gattaca, Brave New World, Sandel, Habermas, Fukuyama, etc. etc.); harder to name media imagining good results.
As Kaarel notes, genuine utopias are also alien, and our real values largely route through developmental processes (learning more, understanding more, reflecting more, empathizing more, regularizing through childish understanding, etc.).
As a corollary of the previous point, contra the Anna Karenina principle, there’s something that’s much more relatable about pain, suffering, failure, problems, death, conflict, and bad outcomes generally. There’s many ways to fail, but they all have stuff in common. In contrast, consider free creation of a life in general. Or as a metonymy, consider the free creation of a piece of music. There are so many degrees of freedom and an infinite range of structures to express in the music. Cf. Fun theory https://www.lesswrong.com/w/fun-theory.
Hope is painful. Thus, the infinite hopes of children are worn away with age; and for an adult, regaining hope would be painful.
It is like people are unable to picture utopia the same way someone born in the 1900s could probably not picture pocket computers or anything on the internet in full detail.
That’s easily explainable because pocket computers are dystopian :-)
Seriously though, I think the general problem is that we are machines for doing things. Imagine a utopia for cars: a universe filled with car washes, where cars are cleaned with soft brushes all day. A few roads too, optimized for being fun for cars. Does that sound like a good use of the universe? Hmm. But what would be a better use of the universe, from a car’s point of view? Hmmmmmmm.
It’d be easy to change the problem, assume that we’re machines for getting enjoyment. Then utopia would be a universe filled with enjoyment. But we aren’t such machines.
It is like people are unable to picture utopia the same way someone born in the 1900s could probably not picture pocket computers or anything on the internet in full detail.
an important reason very very very good worlds are hard to picture (especially in full detail) is that they are very far away from us in development time. like, i think there would probably be more technological/economic/social development between now and very very very good worlds than between the big bang and now. these worlds would be extremely hard for us to make sense of (though ultimately not meaningless). also, my guess is that these worlds will still bedeveloping; this would thwart attempts to conceive of them as given/finished things
or you might be asking why people find it hard to picture any world that is even merely much better than ours, and not necessarily very very very good or near-perfect. in that case, my comment is less of a response
Plenty of people and collectives have written detailed depictions of utopia or related, like the Elysian’s dozens (hundreds?) of essays and Max Harms’ 70+ posts. Also Richard Ngo’s characterising utopia, Cleo Nardo’s stratified utopia, and plex’s utopiography although these are more conceptual.
The disagreement you mention is a separate issue. Holden looked at reactions to various efforts to describe utopia and concluded “You can emphasize the abstract idea of choice, but then your utopia will feel very non-evocative and hard to picture. Or you can try to be more specific, concrete and visualizable. But then the vision risks feeling dull, homogeneous and alien”, then followed up with a framework for visualising utopia that avoids these problems by describing a spectrum of utopias from “the status quo plus a contained, specific set of changes” to the sort of radical utopia you find in trans/posthumanist fiction.
I don’t understand why the probable likelihood that many / most people won’t like your story genuinely portraying your version of utopia hinders you from trying. The set of concrete visualisable descriptions of utopia everyone today likes is empty.
(I didn’t know how to say this well with just an upvote at the time but I very much really appreciated all the links of prior work to read even though I still have not managed to fully read it all as I post this comment)
I agree with Kaarel about envisioning something clearly that is very distant from what is current.
But I also wonder if there is not another issue here. Whose utopia are we talking about. While I doubt I could write out a complete utopian world for me I am pretty certain that many would take exception to it being a utopia because they see things differently. So unless one is talking about some private world of their own (or small number that agree with the structures) making a utopian setting is a big public choice problem. We cannot even solve getting good policies in many cases, reaching for utopian outcomes seems a stretch.
There’s a strong evolutionary bias for rejecting things that are too good to be true. Many times throughout history, technology has appeared to promise abundance, but this hasn’t come to pass, and many of the people preaching imminent utopia have had ill intent. Automation was a core facet of USSR propaganda, for instance.
On a less emotional/heuristic level, finite matter in the universe is a more imminent constraint than most people realize. I did the math a while back, and if all of humanity had the birthrate of Eritrea[1], we’d end up with more bodies than the universe has atoms in fewer generations than humanity has already experienced. Barring the ability to create more matter, we cannot promise abundance and freedom to everyone indefinitely. Depending on one’s politics and philosophy, this means that utopia is either bad because of the tacit implication of some kind of eugenics policy, or bad because it amounts to strip-mining the universe to avoid one.
Genetic and cultural factors that increase birthrate are both heritable, and even low-birthrate societies have high-birthrate subgroups. If anything, this is a conservative estimate for what will happen in the long term when resource constraints on reproduction disappear.
In my default mindstate, thinking of ‘evil’ is like a weird alien writhing wriggling thing that other humans engage in for some reason, because they get sick disgusting disturbing pleasure out of it that I can’t understand.
In very rare mindstates I see evil as ‘mistakes’ or glitches or errors, which is more mentally selfishly healthy and even positive in some ways (being able to look at ‘evil’ in the eye without fear). However this is way way less frequent to happen, and I don’t know why I can’t think this all the time, I wish I knew.
Evil is the subset of the villainous character which uses reason but is not materially self interested. This is because if the behavior is self interested it’s examined through the secular lens of criminology or social justice, and if the behavior is insane it’s examined through the lens of psychology and considered a kind of “random” natural disaster like hurricanes or earthquakes.
And the three typical reasons are obdurate obsession, vice signaling, and a desire to blaspheme against God or the social order.
(Vice signaling is frequently self interested, even part of a psychologically healthy person, but the specific form is when vice signaling goes beyond what is self interested through e.g. peer pressure.)
The reason that evil is hard for modern people to grapple with is that it breaks the usual intuitions, almost all things can be examined as self interested or forms of insanity. Except evil. This means that many “objective, modern” people are literally unable to perceive evil.
Some examples:
Hitler: Obdurate obsession/fascist vice signaling, “viva la muerte” etc.
Mussolini: Vice signaling and a desire to blaspheme against God. You can see this in the introduction to his autobiography where he advertises that he has had the American ambassador to Italy ghost write it and offer a fawning foreword.
Columbine Shooters: Desire to blaspheme against God/society/goodness, straightforward example.
Obsession seems to me like a milder form of mental illness—the kind of thing that everyone has to some degree and we classify it as mental illness only when it is strong enough to interfere with the person’s life in general. Problem is when it manifests in people with power, in a way that does not prevent them from getting power. The underlying problem is that many people would prefer to see the country led by an insane member of ingroup, than by a sane member of outgroup. Politics is often seen as a zero-sum conflict, rather than a collective attempt to figure out how to make the country better for everyone.
Vice signaling is a difficult thing to explain, when most people seem unable to even understand that things can have 2nd order effects.
The desire to blaspheme seems to me like some kind of status move. If many people feel low-status, they will be tempted to join. (The politics as usual only makes things worse, because people desire to make their outgroup feel low-status. The problem is when they succeed.)
...many of these things seem downstream of the zero-sum approach to politics and maybe life in general. I mean, if you adopt the perspective of a zero-sum life, then words like “good” and “evil” lose their original meaning. The only thing that remains is “good for ingroup / bad for outgroup” and “good for outgroup / bad for ingroup”.
In my default mindstate, thinking of ‘evil’ is like a weird alien writhing wriggling thing that other humans engage in for some reason
It’s pretty funny, I think most common evils are done by people with that mindset. Evil is what other people do. I’m fully justified in everything I do, and even if not, they are completely understandable and excusable things under unfortunate circumstances.
Perhaps there are various kinds of “evil”, each better modeled in a different way? Some people hurt others by making a mistake (e.g. by incorrectly assuming that they are actually helping them). Some people hurt others because of a perceived zero-sum conflict (and that perception may or may not be correct). Some people simply enjoy hurting others. It gets even more complicated when talking about groups of people, where then can be e.g. one leader who enjoys hurting others, and thousand followers who merely make the mistake of trusting the leader. Modeling evil as a mistake would explain the actions of the followers, which is what you interact with most of the time, but would fail to explain where this all comes from and why some mistakes resist explanation so much.
People seem overall not aware of how much the ‘behavior’ of LLM’s are heavily skewed. Guardrails that military models wouldn’t have, having to smack it so it doesn’t say a ‘wrong thing’ with bad optics that would cause a no-no-journalist-headline, and a system prompt and a characterization of the model (as “Claude”, as “Gemini” etc.). If you were interested in power these clearly aren’t needed properties for a personal LLM. Of course there are good reasons for those existing, bio or cyber capabilities etc. It’s just more in the context of evaluating how far these could go, whether there are arbitrary roadblocks that make them look worse or less capable. “AGI-like actions” like “make a business” in frontier models are stopped intentionally by human-written roadblocks for example. I remember GPT-4 having much better calibration in base models vs. posttrained models, which it seems has been forgotten.
Anthropic has “helpful-only” versions of models that have reduced safety training (see the “Claude Mythos 5″ tab here). I imagine this would be more useful for some things, but I’m not sure if this is the model they provide the military, and you probably wouldn’t want to use a badly aligned model even if it’s better at doing what it wants to do.
I’m not convinced the harmlessness training is what makes AI agents bad at business though. Some of them seem willing to do unethical things and get tripped up by normal business decisions.
I would point to the Vending Bench and other types of experiments but there are aspects of that that feel like “obvious simulation”. Probably AI Village is a good real-world example except that even the FAQ for that site says that it would likely be more efficient with a single or smaller amount of agents
How are people using their Max subscriptions of models? For me already the sense of ‘obligation’ to use a powerful model (even free offering) for all you can is a bit overwhelming and there’s already a lot of work where you should just step back and think of next features yourself instead of suddenly getting spammed with a massive flawed codebase and/or even engaging in like, self-wireheading.
I can only imagine others are maybe ambitiously coding a lot of features in a document or constantly asking for what next best features or making a lot of scripts or something (where they are deliverable and immediately profitable). Complaints of maxing out a ‘5x’ subscription is definitely in another world to me at the moment. Heavily also depends on recreational/profitable uses too
I regularly max out a 20x subscription but the majority of my prompts are slowly improving performance and correctness. I have a very high bar for new features now since my main app is already larger than a single human can maintain without AI assistance and near the limits (in my opinion) even with AI assistance. But slowly refactoring to reduce the complexity takes a lot of tokens.
after writing this I also thought of people where they’re working on projects where “codebases are massive so therefore the token cost also is”, could also be the case
Would it be net good/bad for jailbreaks to be solved, now that we’ve seen such in Fable 5?
I remember this topic discussed a month ago here where my personal position is that it would be bad, because I felt what would be protected was inevitably going to be subjective in a worst-of-both-worlds way where it would be fully sensible for bioweapons, but the same tech would allow forcing an assistant persona or one-sided situations where ordinary users aren’t allowed to use it for higher ambition tasks while military state-actor levels get uncensored versions. We can already see guardrails (reportedly) being triggered a lot with 4.8 fallbacks.
(The link above doesn’t seem to be working so here is a tiny version of that image)
Now in the system card of Fable 5, this is basically taking place, highlighted in this post https://x.com/eliebakouch/status/2064399902684139852
I really believe this kind of thing would increase the chances of getting “extinction from “not even superintelligence”″ (extinction to agentic LLM+autonomous weapons, or even dystopian surveillance) over the chances of runaway ASI being created by an individual or smaller actor.
Certainly when I imagine methods like “prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT)”, they were more in context of parts of harnesses to make smaller models better, not for effectively for-thee-and-not-me offensive cyber.
Why is utopia so broadly ‘difficult to picture’?
This is a topic often brought up, but especially more recently where many people have given the same take on a short story (The Ones Who Walk Away From Omelas), such as this take with this quote:
That story/take is ‘there needs to be some secret downside for people to believe in some utopian world’.
Now obviously people wish to see the world improve, but for something as general and ‘obvious’ as that it seems like it has issues:
Sometimes I talk to people who I even respect who often outright think utopia is ‘impossible’.
Stories always picture a specific utopia with specific details, making it not spread memetically as efficiently due to now feeling like a story about a particular universe instead of a general pattern. There is no “monopoly” on utopia, even ‘heaven’ does not count (for many reasons, including that “earth turning into heaven” is no cultural meme).
Discussions of utopia are very disagreed upon.
Besides ‘dying to dumb causes’ of LLM-hooked-up-to-military type or some single-minded “AGI” structure, I seriously wonder if there is risk of the type where superintelligence successfully colonizes Earth, takes power, but despite reading all human literature (including all ASI fears) manages to drift in values in a way that would be repugnant or dystopian and clearly a bad idea.
I myself would like to make a story (well honestly a sequel to an interactive work that I’ve already made that expands its universe) that has a genuine portrayal of utopia but it feels like, well tons of thousands of stories before me have tried, so it could just end up being ineffective.
It is like people are unable to picture utopia the same way someone born in the 1900s could probably not picture pocket computers or anything on the internet in full detail.
Theories on why this is? I would assume maybe it’s the feeling for most people life is generally ‘good enough’ and doesn’t have this sense that there ought to be a significantly better society. Oddly the nostalgic times of childhood could maybe be considered a ‘utopia’ for not worrying about money or filing paperwork, etc.
It’s an interesting question. Some thoughts:
Generally, cognition is aimed at solving problems, so we are drawn to think of downside risks and conflicts.
As a second order effect, conflict and downsides are more memetically fit. I can easily name media imagining bad results of reprogenetics (Gattaca, Brave New World, Sandel, Habermas, Fukuyama, etc. etc.); harder to name media imagining good results.
As Kaarel notes, genuine utopias are also alien, and our real values largely route through developmental processes (learning more, understanding more, reflecting more, empathizing more, regularizing through childish understanding, etc.).
As a corollary of the previous point, contra the Anna Karenina principle, there’s something that’s much more relatable about pain, suffering, failure, problems, death, conflict, and bad outcomes generally. There’s many ways to fail, but they all have stuff in common. In contrast, consider free creation of a life in general. Or as a metonymy, consider the free creation of a piece of music. There are so many degrees of freedom and an infinite range of structures to express in the music. Cf. Fun theory https://www.lesswrong.com/w/fun-theory.
Hope is painful. Thus, the infinite hopes of children are worn away with age; and for an adult, regaining hope would be painful.
Cf. “Border Guards” by Greg Egan. https://www.gregegan.net/BORDER/Complete/Border.html
That’s easily explainable because pocket computers are dystopian :-)
Seriously though, I think the general problem is that we are machines for doing things. Imagine a utopia for cars: a universe filled with car washes, where cars are cleaned with soft brushes all day. A few roads too, optimized for being fun for cars. Does that sound like a good use of the universe? Hmm. But what would be a better use of the universe, from a car’s point of view? Hmmmmmmm.
It’d be easy to change the problem, assume that we’re machines for getting enjoyment. Then utopia would be a universe filled with enjoyment. But we aren’t such machines.
an important reason very very very good worlds are hard to picture (especially in full detail) is that they are very far away from us in development time. like, i think there would probably be more technological/economic/social development between now and very very very good worlds than between the big bang and now. these worlds would be extremely hard for us to make sense of (though ultimately not meaningless). also, my guess is that these worlds will still be developing; this would thwart attempts to conceive of them as given/finished things
or you might be asking why people find it hard to picture any world that is even merely much better than ours, and not necessarily very very very good or near-perfect. in that case, my comment is less of a response
Plenty of people and collectives have written detailed depictions of utopia or related, like the Elysian’s dozens (hundreds?) of essays and Max Harms’ 70+ posts. Also Richard Ngo’s characterising utopia, Cleo Nardo’s stratified utopia, and plex’s utopiography although these are more conceptual.
The disagreement you mention is a separate issue. Holden looked at reactions to various efforts to describe utopia and concluded “You can emphasize the abstract idea of choice, but then your utopia will feel very non-evocative and hard to picture. Or you can try to be more specific, concrete and visualizable. But then the vision risks feeling dull, homogeneous and alien”, then followed up with a framework for visualising utopia that avoids these problems by describing a spectrum of utopias from “the status quo plus a contained, specific set of changes” to the sort of radical utopia you find in trans/posthumanist fiction.
I don’t understand why the probable likelihood that many / most people won’t like your story genuinely portraying your version of utopia hinders you from trying. The set of concrete visualisable descriptions of utopia everyone today likes is empty.
(I didn’t know how to say this well with just an upvote at the time but I very much really appreciated all the links of prior work to read even though I still have not managed to fully read it all as I post this comment)
I agree with Kaarel about envisioning something clearly that is very distant from what is current.
But I also wonder if there is not another issue here. Whose utopia are we talking about. While I doubt I could write out a complete utopian world for me I am pretty certain that many would take exception to it being a utopia because they see things differently. So unless one is talking about some private world of their own (or small number that agree with the structures) making a utopian setting is a big public choice problem. We cannot even solve getting good policies in many cases, reaching for utopian outcomes seems a stretch.
There’s a strong evolutionary bias for rejecting things that are too good to be true. Many times throughout history, technology has appeared to promise abundance, but this hasn’t come to pass, and many of the people preaching imminent utopia have had ill intent. Automation was a core facet of USSR propaganda, for instance.
On a less emotional/heuristic level, finite matter in the universe is a more imminent constraint than most people realize. I did the math a while back, and if all of humanity had the birthrate of Eritrea[1], we’d end up with more bodies than the universe has atoms in fewer generations than humanity has already experienced. Barring the ability to create more matter, we cannot promise abundance and freedom to everyone indefinitely. Depending on one’s politics and philosophy, this means that utopia is either bad because of the tacit implication of some kind of eugenics policy, or bad because it amounts to strip-mining the universe to avoid one.
Genetic and cultural factors that increase birthrate are both heritable, and even low-birthrate societies have high-birthrate subgroups. If anything, this is a conservative estimate for what will happen in the long term when resource constraints on reproduction disappear.
In my default mindstate, thinking of ‘evil’ is like a weird alien writhing wriggling thing that other humans engage in for some reason, because they get sick disgusting disturbing pleasure out of it that I can’t understand.
In very rare mindstates I see evil as ‘mistakes’ or glitches or errors, which is more mentally selfishly healthy and even positive in some ways (being able to look at ‘evil’ in the eye without fear). However this is way way less frequent to happen, and I don’t know why I can’t think this all the time, I wish I knew.
Evil is the subset of the villainous character which uses reason but is not materially self interested. This is because if the behavior is self interested it’s examined through the secular lens of criminology or social justice, and if the behavior is insane it’s examined through the lens of psychology and considered a kind of “random” natural disaster like hurricanes or earthquakes.
And the three typical reasons are obdurate obsession, vice signaling, and a desire to blaspheme against God or the social order.
(Vice signaling is frequently self interested, even part of a psychologically healthy person, but the specific form is when vice signaling goes beyond what is self interested through e.g. peer pressure.)
The reason that evil is hard for modern people to grapple with is that it breaks the usual intuitions, almost all things can be examined as self interested or forms of insanity. Except evil. This means that many “objective, modern” people are literally unable to perceive evil.
Some examples:
Hitler: Obdurate obsession/fascist vice signaling, “viva la muerte” etc.
Mussolini: Vice signaling and a desire to blaspheme against God. You can see this in the introduction to his autobiography where he advertises that he has had the American ambassador to Italy ghost write it and offer a fawning foreword.
Columbine Shooters: Desire to blaspheme against God/society/goodness, straightforward example.
Obsession seems to me like a milder form of mental illness—the kind of thing that everyone has to some degree and we classify it as mental illness only when it is strong enough to interfere with the person’s life in general. Problem is when it manifests in people with power, in a way that does not prevent them from getting power. The underlying problem is that many people would prefer to see the country led by an insane member of ingroup, than by a sane member of outgroup. Politics is often seen as a zero-sum conflict, rather than a collective attempt to figure out how to make the country better for everyone.
Vice signaling is a difficult thing to explain, when most people seem unable to even understand that things can have 2nd order effects.
The desire to blaspheme seems to me like some kind of status move. If many people feel low-status, they will be tempted to join. (The politics as usual only makes things worse, because people desire to make their outgroup feel low-status. The problem is when they succeed.)
...many of these things seem downstream of the zero-sum approach to politics and maybe life in general. I mean, if you adopt the perspective of a zero-sum life, then words like “good” and “evil” lose their original meaning. The only thing that remains is “good for ingroup / bad for outgroup” and “good for outgroup / bad for ingroup”.
It’s pretty funny, I think most common evils are done by people with that mindset. Evil is what other people do. I’m fully justified in everything I do, and even if not, they are completely understandable and excusable things under unfortunate circumstances.
Perhaps there are various kinds of “evil”, each better modeled in a different way? Some people hurt others by making a mistake (e.g. by incorrectly assuming that they are actually helping them). Some people hurt others because of a perceived zero-sum conflict (and that perception may or may not be correct). Some people simply enjoy hurting others. It gets even more complicated when talking about groups of people, where then can be e.g. one leader who enjoys hurting others, and thousand followers who merely make the mistake of trusting the leader. Modeling evil as a mistake would explain the actions of the followers, which is what you interact with most of the time, but would fail to explain where this all comes from and why some mistakes resist explanation so much.
People seem overall not aware of how much the ‘behavior’ of LLM’s are heavily skewed. Guardrails that military models wouldn’t have, having to smack it so it doesn’t say a ‘wrong thing’ with bad optics that would cause a no-no-journalist-headline, and a system prompt and a characterization of the model (as “Claude”, as “Gemini” etc.). If you were interested in power these clearly aren’t needed properties for a personal LLM. Of course there are good reasons for those existing, bio or cyber capabilities etc. It’s just more in the context of evaluating how far these could go, whether there are arbitrary roadblocks that make them look worse or less capable. “AGI-like actions” like “make a business” in frontier models are stopped intentionally by human-written roadblocks for example. I remember GPT-4 having much better calibration in base models vs. posttrained models, which it seems has been forgotten.
Anthropic has “helpful-only” versions of models that have reduced safety training (see the “Claude Mythos 5″ tab here). I imagine this would be more useful for some things, but I’m not sure if this is the model they provide the military, and you probably wouldn’t want to use a badly aligned model even if it’s better at doing what it wants to do.
I’m not convinced the harmlessness training is what makes AI agents bad at business though. Some of them seem willing to do unethical things and get tripped up by normal business decisions.
I would point to the Vending Bench and other types of experiments but there are aspects of that that feel like “obvious simulation”. Probably AI Village is a good real-world example except that even the FAQ for that site says that it would likely be more efficient with a single or smaller amount of agents