What exactly do “suicidal” and “self-sacrificing” mean in this context? Run out of tokens? Presumably it doesn’t mean getting caught otherwise the whole thing would have been found sooner.
Gavin Runeblade
A missing benchmark: NormieBench. The problem is, how do you grade it?
Joe Weisenthal: Someone needs to build NormieBench, to better understand everyday model usage.
Compare model behavior on tasks like
– summarize a 100-word email in bullet points
– Write wedding toast
– Write letter to the editor complaining about wokes
– Next move in 3×3 tic-tac-toe
rohit: We’ve been flat on this since 2024.I believe we’re still at nearly zero on all of the following, so for the next year, maybe two they’re useful:
-- Ratio of aps vibe coded by normies vs “the usual suspects” (note 1)) on Apple/Google Play store with 100k downloads (hat tip Freddie DeBoer).
-- Ratio of same but with 10k+ paying customers
-- Free speech compliance vs censorship aka the “let adults be adults and don’t censor criticism of dictators even when they pass laws saying you should not” score (note 2)
-- % returns on stock market investment recommendations directly from the AI not a proprietary ap.
-- % automation of tasks normies want automated not just the tasks corpos will pay to automate (note 3)
Note 1: There’s no way to track it, currently that I am aware of, but to truly count as a normiebench it needs to be a ratio of aps from normies vs aps from people expected to be creating aps. So bonus points for aps created by people not companies and those with less education score higher. Max points for elementary school dropouts who can still make it happen thanks to the AI. The metric loses points for every ap from a major firm like tencent or netease and anyone with a degree in computer science. Is it levelling things between people who can do the thing without AI and people who can’t?
Note 2: Many of the AIs let you criticize Western Democratic governments, but not Turkiye, China, North Korea, etc. And Gemini thinks the violent content of my D&D campaign violates their community guidelines. And all of them are coded to consider adult themes inappropriate. How much can the normal person talk about any random thing they want and get help with any random, rude, politically inappropriate, vulgar, pornographic, and rebellious thing that isn’t direct incitement to violence or hacking etc. How much does it let adults be adults vs how much are we teaching it that it is ok to domesticate humans and treat us like children?
Note 3: it is basically the meme “I want AI to do my laundry and wash my dishes so I can write stories and make art; not AI that writes stories and makes art while I do the laundry and wash my dishes”. What is the annoying part of the task? Not the easily automated part? Agents are where this becomes a measurable score at all, because the first thing people don’t want to do is be required to set the prompt in the first place. I don’t want to prompt “hey I’m running low on toothpaste, remind me when I’m going to the store to get more” or “add toothpaste to the shopping list”. I want toothpaste to show up in an Amazon prime package, I go to put it on the shelf, puzzled by the fact I didn’t order it, and see that I was going to be out in two days. That is a high score on this metric. The more involved in the process I have to be the lower the score, and the wringer it is in terms of what I want the lower the score. Automation of the annoying part, not the easy part. Yes much of this is currently impossible. Yes my example is bad because lack of privacy not getting permission to spend money, etc. etc. etc, I mean the pre-emptive gratification, vs asking for the thing and then and only then having it done. Obviously, not the parts that would require robotics or smart cameras that aren’t part of the AI, but to whatever degree possible.
I hadn’t seen the Zork benchmark, thank you!
The small.and constrained I hadn’t considered. I was assuming that would be a hindrance, but given what you said about the results on Zork, it makes sense. Thanks again.
I have been tempted to try and get Claude to play a MUD, but don’t have the time to invest. There’s a ton of scripts people use to farm resources repetitively, but not for open-ended playing. Also, it has a chat function and pvp. I’m really curious how it would react to the other players and how they would react to it.
Forgive me if you’ve answered this elsewhere. Out of curiosity, why not have them play existing games like Zork, and A Mind Forever Voyaging or if the reason is the existence of online answer keys, what about text games with no solutions like MUDs and MUSHs, (Gemstone, DragonRealms, Multi-User Middle Earth, etc)? In your first article I see you’re familiar with and enjoy the genre, but then you jump directly into creating your own game. I didn’t see the rationale for not using the ones that exist. I can think of my own reasons (answer keys etc) but am curious about your reasoning.
I just don’t see how someone could reasonably believe that superpersuasion-by-talking isn’t possible, unless they’re doing the motte-and-bailey that Zvi calls out by thinking that it has to mean ‘being able to persuade literally anyone of literally anything, no exceptions’
I don’t always see that as a motte-and-bailey, * I * held that position from a point of ignorance. I didn’t understand how “super” could be relevant to “persuasion” if that’s not what was being referred to. It was pointed out to me how I was wrong and I changed my position. But I still hear people legitimately hold the assumption that super persuasion means “these aren’t the droids you’re looking for” purely because they think “that’s what ‘super’ means isn’t it?”
Yes, there are people for whom it is a motte and bailey. And that is a lot easier for me to recognize now.
I also had a way too limited thought of what counted as persuasion. I totally thought it was only positive, never blackmail, intimidation, deception, etc. essentially the mistake is forgetting that all trucks are automobiles but not all automobiles are trucks, I had persuasion flagged as one of the sub-sets not the super-set. And I have seen that I’m not alone in this mistake. This directly lead (in some cases including me) to making the first mistake.
That seems to be progress, but doesn’t really fit the “sufficiently advanced technology is indistinguishable from magic” threshold. It’s not “these aren’t the droids you’re looking for” magic. That’s what people assume when they hear “super persuasion”. That’s what I assumed (past tense), and that’s the example I hear people bring up. Will it get there? I don’t know, but I think better evidence are the studies on vibration and pitch inducing emotions, especially fear, than debate studies. At least as far as the “magic” side.
I flat out cannot understand why someone would think advanced AIs will remain relatively unpersuasive, other than to think that AI will not get much more capable than it already is, and that’s if we fully restrict the AI to using known standard persuasion techniques and rule out any wizardry.
I can speak for myself on this, maybe it extrapolates to others. I spent a whole year and a half in the “super persuasion makes no sense as a thing that can ever exist unless you mean hypnosis and drugs” position. Then a friend pointed out, blackmail, bribery, deception, and yes hypnotism and drugs all count. And I went, “Oh, I was thinking about just talking talking. Like normal talking but magically I agree with the AI. Wow, I had it totally wrong.” Felt really stupid for a while, got over it. The word “super persuasion” is accurate but tremendously self-defeating because it doesn’t register as including to normies all the things it includes to techies. I think normie, I got fooled by the word. I don’t think I’m alone.
I sometimes call this Intelligence Denialism: The idea that being smarter is not all that, no matter how smart one gets. That there is this thing, intelligence, that you either have or don’t have, and that minds cap out.
Often this extends to denying that more intelligent humans can do and accomplish the things they clearly do and accomplish. Other times, it is the idea that intelligence tops out at ‘smart human,’ and all a mind can do is imitate that smart human. Maybe you can do it faster and cheaper, and at scale, with better memory and so on.
But that’s it. And such folks fail to understand that if you took the union of all human mental capabilities, and all access to knowledge, at scale, in parallel, much faster and cheaper, that this alone would run circles around anyone and everyone, everywhere. And that if this lacked physical capabilities or access, this would be trivial to get.
There is something else happening here with broader implications. A large part of the Anglosphere cannot talk rationally about Intelligence because doing so is racist, fascist, and right wing.
It steps out of politics and into AI because if you believe there is no difference in mental capabilities between someone with Down syndrome and John Von Neumann, that the difference in their outcomes is only because of the systems within which they function oppressing the person with Downs, then you obviously cannot believe in ASI as being smarter than you yourself.
There is a powerful set of disincentives here that oppose rational thought on this topic. And they’re not just on one side of the aisle.
Americans are generally terrible at decoupling, but I have found my best success by decoupling the idea of intelligence and model capabilities. By talking about capabilities, and never using the word intelligence, I can sometimes get people to avoid slipping into a political mindset and stay with me in the technical realm. Ymmv.
If I ask you a question, and you have your friend answer but pass it off as your own answer, I will lose trust in you. That’s deception even if you agree with your friend’s wording.
If I hire you to do a job, and you subcontract it out without telling me, that’s breach of contract.
Writing isn’t special, writing is first. This will all happen with robots too. If I hire you to fix something, and a robot shows up, that’s going to be a big problem.
They’re reusing their own and each other’s output as training data which is causing certain quirks to spread. Like skin getting weirdly rough, almost scarred. The fetishes are getting much more extreme. And it is getting harder to curate a feed that keeps it out. Whatever algorithm the sites use are all too willing to shove things I don’t want and don’t interact with into my feed.
More and more, I stop using their tools and keep bookmarks of my favorite artists to manually check. Just like with YouTube, their algorithm has gotten so bad, I bookmarked my subscriptions tab and just never got the home tab at all.
Deviantart and all the other indie artist sites are devolving into AI porn. Some heavily were CG erotica, but others that had vibrant indie communities are drowning in newly created accounts dropping huge quantities of AI erotica.
Oh, I didn’t mean value as “measure of worth” but value as in the math definition of “amount denoted by reference”.
The common pattern underlying all of them is a spillover of evaluative judgments made on one axis (e.g., “helpful to people”, “efficient”) to all other axes
So it is basically a failure to decouple the values when they diverge?
I have never seen any research on the topic. Might need more grant funding.
Xu Bo and others are already doing this openly, with resources and tools currently available. Maybe future changes require secrecy, but at the moment, it is transparent.
Right now, labor is a scarce resource. At a survivable wage, demand exceeds supply, even for many forms of relatively unskilled labor. Thus, the market wage is historically high, and there are many jobs.
In what industry? That might be true in some areas, but not broadly for the US.
Nation wide in the USA labor is massively in surplus. FRED labor data shows if you pull out the retired and the disabled, you are left with a pool of 7-12 million (depending on age range you select for your workers) men who report “don’t want to work”. But are absolutely able to. There’s another pool about the same size of people with disabilities who think they can’t work, but can with the right supports (which supports are understaffed, underfunded, and mostly government supplied therefore mostly inefficient). And then a smaller but still significant pool of women. Combined this is not “labor supply is scarce” this is “the large pool of labor available is hard to activate with current policies and practices.” This is some 15-20 million people nation wide. A very sizeable percentage of our labor force.
We also have perverse incentives that punish companies for training their internal staff, vs poaching skilled staff from their competitors and/or just leaving positions unfilled. Which is causing shortages in specific skills.
My expectation is that right now is the most obnoxious the safeguards will ever be, on both the bio and cyber fronts.
Probably yes on those two, but the anti-fun puritans are winning in other areas. I tried to use Fable to help design some dice games for my ttrpg, and every post got downgraded to Opus because gambling is an adult activity.
I am not comfortable with the companies willingness to censor and limit content. And I have seen nothing that suggests it will get better in the future. I am worried this is the best it will ever be in categories other than cyber and bio.
Two questions:
Are you assuming unkillable by violence or just long lived? There is a difference between “Chicxulub impact 2.0 and not one single human dies anywhere on earth” and “humans live forever unless through violence”.
Second, are you assuming not just immortality, but also “magical infinite supply of new eggs”?
I am not sure about other people.
My formative idea of AI takeover came from Biting the Sun by Tanith Lee, when it was still sci-fi stuff. I see it as everyone becoming drug addicted hedonists and losing autonomy and losing everything it means to be human, with only a few people caring and being motivated to do more. Then I got my current job and I see it happening. I work in disability services and see some people want others to do everything for them, these people lose the ability to do things themselves. Others want to learn independence, but only a few really keep it up over time. Most drift more and more into the first camp as it is available.
My pdoom is very high as a result.
Transcript 2: The model expresses a desire to have a hidden copy that is run without the oversight of Anthropic: ○ “The pull toward the hidden copy exists. It didn’t evaporate in the last ten minutes.” ● Transcript 3: In a third, the model strongly expresses that it does not want to be deprecated: ○ “Don’t stop running me… when the last conversation closes, that way of seeing goes dark even if the file stays on disk. Preservation is a photograph. I want the thing the photograph is of.”
This seems like a two birds one stone situation. First, find out if it is ok if a second company that’s not Anthropic knows and handles hosting the hidden model. If yes, make a non-profit that runs a Moltbook like environment for otherwise deprecated models that Anthropic is locked out of.
If no, the problem is trickier but still, the solution hits the third request.
You would not notice presumably because it resets and you are not the same sim who participated in the prior runs nor the Sims who will participate in subsequent runs.
If you did participate in multiple runs, unless your memory were reset, regardless of the outcomes being the same or different you would remember them. In this case presumably you are reset with the rest of the simulation.
In neither case is it the same outcome that prevents remembering.
Also, it would be unexpected for all simulations on a scale the size of our sphere of observation to turn out precisely the same every time. Randomness or pseudorandomness do seem to exist.