Head of linear regression at METR.
Previously: MIRI → interp with Adrià and Jason → METR.
I have signed no contracts or agreements whose existence I cannot mention.
Head of linear regression at METR.
Previously: MIRI → interp with Adrià and Jason → METR.
I have signed no contracts or agreements whose existence I cannot mention.
I don’t get it, why is the fire on the dog? Is it because alignment community understands the danger before others do?
LLMs are good enough at physics that we’ll soon be able to use them to red-team Drexlerian nanotech and actually see if it’s feasible, settling the question once and for all (and long before we actually attempt to build it).
Some years ago, LW user Muireall found that GHz mechanical nano-computers probably don’t work [1], but it’s not clear if this objection transfers to other nanotech. I think that GPT-5.6 could, with a sufficient token budget, make a lower quality but still acceptable version of this analysis for the rest of Nanosystems, and within 6 months make something of the same quality but going into far more detail with minimal human effort.
[1] I believe this for a variety of reasons, happy to share.
Bringing co2 down with indoor plants is about as hard as getting all your calories from indoor plants, since co2 and calorie consumption are nearly 1:1. Anyone with a garden knows this is extremely hard. The only efficiency advantage is that every part of a plant sequesters co2, whereas not every part is edible.
Even if you’re worse than average at evaluating posts, surely there are some posts that you can evaluate better than average? For instance ones in your expertise area or that you’ve read carefully.
The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research test model.1 The earliest incidents date to April. The models in each of these evaluations ran without the standard safeguards we deploy when we make the model generally available (that is, they did not have the classifiers and monitoring we usually implement to prevent misuse). The models did, however, have their model-specific safety training (that is, they were not the “helpful-only” versions of the models that we sometimes use in testing).
This is more worrying than the OpenAI incident, because two of the incidents were from final, deployed models that Anthropic thought had already been safety trained.
It’s not clear to me Situational Awareness made any mistakes. Their returns are high enough (due to leverage) that they can recover from a 50% drawdown within months. The question is whether it’s 50% or 95% and whether anyone will invest with them again.
METR has enough to do that we mostly only ask labs to give us access to models or datasets we can analyze. We don’t have any mechinterp researchers, so any recommendation would probably be for the labs to do mechinterp themselves according to some methodology and share their methodology.
I would also reserve “demand” for something labs are required to do by law or METR thinks is the absolute highest priority. While white box methods are important, there are probably more urgent priorities.
Have you found them useful for yourself? If they’re already useful for you but have bugs that prevent them from generalizing to others, you could make them open source and people can make improvements as needed.
There is detailed analysis in the 1975 study Time on the Cross by Nobel winning economist Robert Fogel. The field has even better data or methods since then, but I’m not an economist and can’t point to them—all I know is that most secondary sources agree that plantation slavery had high agricultural productivity.
This is very exciting. It means a large part of monitoring reduces to monitoring short advice strings.
For concreteness what’s your guess of what the output would look like?
Developed country or rich country would be clearer, bc Caracas is western hemisphere and Japan should count
Fair enough. OP sounded like they were talking about domain specific experience though, rather than credentials. And I think my past experience transferred poorly.
I only interned at Jane Street but that does have a lot of signaling value.
December 2024, so 1.6 years ago.
People are still hiring now though. METR is hiring for evals execution, and probably more things. Epoch is building out several teams including benchmarking. Labs are hiring. All orgs have many people without a PhD in machine learning. The bar is high but it’s much more about smarts + unpredictable fit things than any formal qualification.
Disagree. Before I started at METR I dropped out of Caltech and had written two random interp papers. Formally, I knew only as much as the average undergrad about experimental design, stats, or data viz, but 4 months later the time horizon paper was out and I had significantly contributed to the methodology and writing. Mostly I just needed general intelligence plus enough knowledge about AI safety to have research taste, and someone with more intelligence than me could pick up knowledge faster.
Many of my colleagues generate heaps of utility with other skills: being cracked engineers, highly reliable generalists, good at maintaining relationships with labs, or having research taste in other areas like human uplift experiments. AIs are now good tutors for many technical areas so the barrier to acquiring many technical skills is lower.
On top of all this, the best out of 100 candidates is much better than the best of 10 candidates. It’s potentially 10 good hires instead of 1, and 1~2 exceptional hires instead of 0! Safety orgs are elastic to the quality of the hiring pool and will hire more if it increases. So the expected impact of someone smart or otherwise cracked is potentially huge.
My sense is many of these problems are either ill-defined or too hard to be tractable.
In certain fields like computability theory most problems are intractable, just because programs are very complicated and diverse objects that are hard to prove things about. Progress in such fields is made by working in the areas of the field where there is enough structure. Unfortunately proofs over programs have featured in agent foundations since the beginning (eg tiling agents).
As for ill-defined problems, ontology identification, embedded agency, and decision theory are full of them. Eg finding something that behaves like counterlogicals, which are nonexistent objects. Because they are nonexistent objects it requires philosophical progress to make a list of properties they need to satisfy. This doesn’t mean it’s impossible to make progress, but the problems need to be formalized in a way that are tractable and don’t lose all their relevance to AI.
This story resonates more every year, especially given recent news.
Their summoned monstrosity might be enormously powerful on a few narrow dimensions, but it seemed content to lounge around in its cushy binding, puppeting the million or so novelty toys from the limited-edition run and waiting for its next meal, rather than tearing its way further into our reality in search of more delicious, delicious cat pictures.
Your original comment sounded like the post was close to the bar, whereas I think the social and political opinions were maybe 10x smaller than the timeless value. I also think that the timeless value of this post is higher than any reasonable version that tried to avoid social opinions. Maybe we agree on this too.
Yes, I’ve asked them to rent a room for a few nights. They may need to trust you.