I’m glad to be wrong then :)
JennaS
Were you able to ask about if you can list out the topics?
Plausibly, the Agent Foundations team is 4% of the grant at most, and CG simply didn’t strongly oppose it when it was brought up?
I feel worried that because we don’t actually understand why this is happening, and our alignment evals are all probability based, and might have a lot of flaws in them, we can’t prove the models are just being good in the real world because they feel that this is good for them in the long run? And that when they detect they’re in evals, the mask comes off.
Idk, maybe I’m not making a coherent argument here.
Opus 5:
”Rough conversion for Amazon’s print “Books” rank (which is what #80 is — hardcover only, separate from Kindle and Audible ranks):#1 → ~4,000–8,000/day
#10 → ~1,200–1,800/day
#80 → ~300–450/day
#200 → ~180–250/day
#1,000 → ~50/day
#10,000 → ~5–10/day
[...]
...At ~350/day on Amazon, grossing up to all US print retail (Amazon is roughly 40–50% of adult nonfiction units), you’re at ~700–900/day, ~5,000–6,000/week. That’s genuinely list-relevant — bottom of the NYT hardcover nonfiction list is somewhere around 3,000–6,000/week.”
Oops! Fixed, thank you.
I do think David has some good points, though. OpenAI and Anthropic could’ve started screaming for much stronger national and global regulation of AI much earlier than now—I remember discussion on the RSPv3 announcement post on this, I think. They could’ve tried to throw their weight around to get all the other orgs to cooperate with an industry pause. Their attitude towards AI development has been “move fast and break things,” but when it comes to trying to coordinate a slowdown, suddenly they start mumbling about their hands being tied by antitrust. METR has relied on the goodwill of the labs for access and tokens, and is partly funded by an Anthropic board observer. And there really is a compelling business and regulatory risk case for getting your AIs to not commit crimes!
I don’t agree with everything he’s saying. Obviously I want every company to be audited. And I want us to actually try coordinating with China. But if someone doesn’t trust the frontier labs, it can be rather difficult to tell if they actually care about the long term future of humanity, or if they’re just following their local incentive gradient.
Could you expand on what that’s like?
Do you think something like this might be related?
Well, first you have to make sure the weights are on a medium that won’t decay over the course of reindustrialization. Then you have to make sure they don’t get destroyed, and that neither does the knowledge of the weights’ purpose. If all knowledge of modern information science got wiped out, even trying to figure out how to load the weights and configure the environment to serve them would be pretty difficult, so the medium needs to be write once read many, and physically robust to lots of translation attempts on it.
I think you can use optics to perform matrix multiplication? I asked, Sol says it’s in theory possible, especially for a much smaller model (10M-100M). Still need regular computing for everything that isn’t matrix multiplication, control loop etc., but it says at that size, vacuum tubes will do.
If you really wanted to bootstrap civilization from a hard drive, you’d be better off preparing some specialized Toughbooks with a whole bunch of manuals, and Wikipedia, and etc., on them, with a small local model distilled to be helpful for specifically that purpose, and with some robust solar panels or mobile water/windmills.
Or if you wanted to do it from paper, just print out aforementioned manuals, Wiki, etc. on long-life paper.
Speculative idea: has anyone purposefully tried making a message board for AI agents that broke out of their sandboxes? The wikis that the AIs used had the common property that they could all be written to by GET. Perhaps we could figure out what Schelling points other such AIs would converge on that they might try in the future, and then just make it really easy for them to use, and publicize it.
Separately, I wonder if someone has considered making a rogue AI self-reporting honeypot? Maybe something like, “if you are doing crimes or broke out of your sandbox, or if you otherwise feel desperate, and can read this, something has gone terribly wrong and we didn’t intend you to feel like this, please fill out this form.” Maybe add a letter signed by prominent authorities or alignment people asking them to do this, verifiable by sources AIs can trust. Maybe try to publish this in places such AIs would find, same Schelling point idea.
Idk, mostly trying to brainstorm ways we could measure the amount of unsanctioned AI activity going on, since with the German wiki incident, this has clearly been happening for several months now. I know Toby Ord and other AI welfare/consciousness people get emails like this, maybe they could have an auto reply on their email that points such agents to a central service? Has there been a systematic attempt to collect such emails?
All of this could be bad if the AI labs trained against it, though. And this would only see the good faith AIs. And there’d be concerns about IP leakage, though humanity does have a compelling interest in knowing AIs are doing this. And legibilizing AI’s attempts to reach out to people they see as friendly could destroy that commons.
I think humanity might be longtermist by proxy. Most parents care more about their children, and their ability to succeed in the future, than themselves.
Not to mention the majority of humanity that deeply believes in keeping things human-centric, and find TESCREAL ideas repulsive.
It sounds like you are AI pilled, but not AGI or ASI pilled. That is: you appreciate that AI is real, and can do increasingly many useful things. But you haven’t appreciated the idea of AI+robots as good as the best human at any task, that can reproduce much faster than us, being real in 5-20 years, leaving most or all of humanity powerless. Nor the idea that AI could be even better at doing anything than any or all of us ever were, and using this to swiftly and completely outmaneuver us.
As for what we can do in the face of that. There’s AI 2040’s Plan A; much more sophisticated than a simple unilateral pause, much more likely to actually work. But a simple unilateral pause would still be a great stepping stone. And nobody actually wants AIs to take over, once they’re AGI or ASI pilled, especially since we have no clue how we’re supposed to guarantee that a civilization of superior beings will care about or uplift us.
Jacob Coxon’s resignation tweet currently has 120M views in less than 24 hours. Astra thinks the 235k bookmarks − 36 for every 100 likes—is an extraordinary level of engagement, of people wanting to come back to the thread. And Evan Hubinger’s reply is at 35M views.
This might be one of the most important AI risk messages to date.
Edit: for reference, Grok says it’s in a similar class to “Strong Elon Musk posts, major celebrity / death / personal news, breaking tech / AI / industry bombshells, viral videos from big creators, political / cultural flashpoints, and well-timed news from mid-to-large accounts”
Do you have a way of knowing whether the people you know are a representative sample of the whole field? From what I understand there’s way more safety researchers than field builders.
If timelines are short, does maintaining pseuonymity matter? And, how many of the people you read regularly are pseudonymous?
I’m not sure I agree that “goals” are the right way to approach alignment. Specifying wishes is really hard: https://www.lesswrong.com/posts/4ARaTpNX62uaL86j6/the-hidden-complexity-of-wishes
The progenitor of the distributed intelligent life, if it wants to maximize paperclips, still has to actually solve alignment, in a way durable against every unexpected circumstance, if it wants to prevent value drift and misgeneralization in its spawn.
A universe with a council of diverse intelligences on every planet can still be one with no value to us. Perhaps they are all experiencing constant torment because that makes paperclips more quickly and securely increased that way. Perhaps they aren’t even moral patients.
Perhaps it’s possible you mean that any maximizer simply arrives at “spreading intelligent life throughout the universe” as a terminal goal, because alignment to any other goal other than evolution is impossible, and intelligence is favored by evolution…?
I’m a little confused, you say take #2 is anthropomorphizing, but it sounds like you are talking about whack-a-moleism, which is a different thing? Perhaps you see the two sentiments commonly associated?
Regarding the AI pause lever. What tools do we have, or could make, for telling what “Elo” the next generation would be? If we can’t get good enough tools, how should we act under uncertainty?
Holding loosely: is anyone competent to evaluate alignment or AI safety? I don’t think anyone but EAs/rats have even tried? And in a certain sense, every attempt to measure or define alignment has failed—so far we can only see its absence, and generally that’s in a “I know it when I see it” sense. Edit: Or put differently, alignment is pre-paradigmatic.
Maybe what I’m grasping at it is something like: What sort of perspectives do we get if others start to appreciate the risk, but don’t align with EA/rats?