DMs open.
Cleo Nardo
The dots separate the first group (which were shared around a few slack meme channels) from the second group (which I added before sharing publicly).
I don’t think this is a big component of the explanation.
Alex Turner: I wasted thousands of hours proving my P theorems. (I still believe P but for different reasons.)
Arguments for P
Yeah I think if someone was truly evaluating projects on INT then this is a decent case for bumping T slightly. In fact, my guess is EAs are already overweighting T because people care about career capital and status. So correcting down on T would increase their impact.
Moreover, I don’t understand the impact case for individuals focusing on their own career capital versus something like “[people with my values] capital”. Like, it’s important for the EA/AIS community overall to have big legible wins, but maybe this implies individual EAs should be focused on the 1-10% range.
“Replaceable, wishes he wasn’t” vs “irreplaceable, wishes he wasn’t” 🙇♂️
Maybe you should do things with a lower probability of success. I think most people doing projects which have a 50%+ chance of succeeding, which is probably a good idea for your career and status. But it might be easier to farm EV in the 1-10% range. This is all very abstract so I’m not sure this is helpful advice to anyone.
(Here’s smth I wrote a month ago.)
Plan A might be a kinda rough situation, given you’d have millions of people lobbying for the project to go faster. Like, 30% of the pop (60-89) is taking an additional 30% chance of death if the singularity happens after 20 years compared with 10 years.
One solution is to cryo everyone 60+. I’m not entirely dissuaded by “cryo is too sci fi for them” because the original worry is that the cohort is lobbying for an early singularity! Maybe I should invest in some cryo startups. I bet even Leopold hasn’t backchained this far.
A similar solution is a “soft cryo”, i.e. medically induced comas. It seems pretty plausible that during an intelligence explosion, doctors will say “Keep that patient on life support for another year, the medicine will be better.”
Another solution is simplying paying risk premiums to the 60-89 age range, to compensate them / stop them lobbying for early foom.
AI has recently solved a number of conjectures. I haven’t checked this, but it seems that these are almost all resolved false, via the construction of counterexamples.
I had the same problem in the UK. I bought them instead online, at www.visiondirect.co.uk/ which doesn’t need prescription. You’ll need to tell the website which contact lens prescription you want to buy, but Claude can calculate that from your glasses prescription. Then practice at home, and you’ll quickly upskill.
I spent ~2 hours trying before I had put both contacts in and taken both of them out. This time was split over like two sitting.
Much of Damon Binder’s work is based on a methodology of speculative engineering, i.e. thinking about the feasibility of various ways to solve a problem. See:
How much this kind of work update [smart, reasonable] people on the questions he discusses?
How reliable is this methodology? I’m not sure.
It seems IMO easy for this methodology to conclude “X can’t happen because I enumerated the ways you’d do X and none of them work” and then X does happen because you missed something.
Or, “X can happen because here’s a method which works” but then X can’t happen because the method actually doesn’t work.
Maybe this is a skill issue, i.e. if I was as smart as Damon Binder I could see that, yes, that really is an exhaustive enumeration of the ways X could happen. And yes, that really is the feasibility of each method.
What’s the best/worst examples of this methodology?
I’m interested in cases where speculative engineering concluded X is feasible (at some point in the tech tree) but it wasn’t — or X isn’t feasible but it was.
What is the relation between Damon Binder and the God of Straight Lines? Where do they disagree? Which diety should I trust more?
I don’t think this proposal should be high priority for labs, compared to (a) ensuring that your monitoring system doesn’t have more direct gaps/holes, and (b) making sure your monitors actually catch the things you’re worried about in realistic settings.
Your proposal might be high priority because it has more bang for your buck because the other things seem much harder. Maybe I’m underestimating how hard this proposal would be to implement.
Yeah, I meant Carl Shulman. Well-spotted, thanks.
Yep, one hope is “EAs are good at forecasting how things will go and that maintains our influence. Both because we have a reputation as good forecasters so people listen to us, and because we make good decisions based on those forecasts (i.e. a similar arbitrage play as the “AI will be a bigger deal than people think, let’s pile into it”)”.
But I’m worried that [the top few] EAs are approaching the Horizon Point beyond which their impressive forecasting breaks down. Like, 10 years ago, Carl Schulman and Paul Christiano [and people in that reference class] could see in such clarity how things would go in 10 years time. But today, they can’t see another 10 years into the future. Like, maybe a few years into the future? Even there I’m sceptical. They might have run out of “alpha”.
For a concrete example, Daniel’s “what 2026 looks like” was pretty on-the-nose. But AI 2027 isn’t nearly so impressive — the distributions are so much broader, and Daniel himself is like “here’s a bunch of ways things could go.” AI 2026 didn’t have multiple scenarios!
One problem here is that our alpha was predicting the change in the technical landscape. But it’s not just technical landscape that will change. It‘s also the political landscape, which has been dormant for most of 2019-2026. And soon the geopolitical landscape will start rumbling. And the economic. And the cultural. I think these things are still pretty dormant, compared to how much they will start flipping the overall strategic landscape. (See here.)
IIUC, this is why Daniel’s current forecasts are choose-your-adventure, compared with his 2021 forecasts about 2021-2026.
Like, do we have good forecasts now which weren’t ~ “priced in” 5 years ago among Carl/Paul/Kokotajlo/etc? I feel like the uncertainties they had 5 years ago are pretty much still unresolved.
iirc, gpt-4-base was handed out liberally to outsiders. Most outsiders didn’t care about it, because it was so much less useful than the post-trained model, but they could’ve got the api keys by sending a slack message.
The models that labs don’t liberally hand out are the helpful-only post-trained models.
When we say METR have “access” to GPT-5.6 Sol — that might mean access to study the models, or access to use the models. In this post, I’m mostly focusing on using the models to accelerate their own R&D.
Could you talk publicly about special model access at Apollo? My impression was no.
I do think that “labs don’t want other orgs to start badgering them for special access” is actually a convincing reason for METR to not disclose any special model access. I want third-party model access to be as cheap and riskless as possible for the labs! At least, in the current regime where we are relying solely on their goodwill.
EAs will have increasingly diminished leverage over the labs. This is mostly because more groups are waking up to ASI, and will start pushing the labs towards their values and worldview.
We might look back on 2021-2025 as a relatively nice regime where the labs only had to please the investors and the EAs. And in 2025-2030 they’ll need to please the investors, the EAs, the natsec, the judiciary, each major government, each major religion, the cultural elite, each industry of workers, the AIs themselves, etc.
Memes do actual epistemic work (e.g. this post, I claim) and on the margin I‘d prefer more memes (in absolute, not proportional, terms). But some cultures only incentivise memes, which would suck for this community bc much of our work is by necessity unfunny.
Fwiw I think our community rewards unfunny stuff, so I’m not too worried. For example, the top LW post is Alex Turner’s which has zero jokes.