Director Detection at SecureBio in Boston. Speaking for myself unless I say otherwise.
jefftk
State of Pandemic Early Warning
Initial DIY Cleanroom Experimentation
Why I Stay Off Twitter
I use this all the time for cooking unfamiliar things, and occasionally for DIY repairs. I haven’t tried it for cleaning the bathroom, but I’d predict it would do well.
neither can they make coffee
That’s quite far from a frontier model:
I tried this with Claude Sonnet 4.5, because it’s free and already available from my phone.
Fable or Astra would have no trouble with this.
I don’t think frontier LLMs can even beat classic pokemon games using only raw in-game images as input.
That’s out of date:
Fable 5 … also needs less scaffolding: for example, previous Claude models struggled to play Pokémon FireRed even with harnesses that gave them additional helpful tools, but Fable 5 beat FireRed with a minimal, vision-only harness.
https://www.anthropic.com/news/claude-fable-5-mythos-5
Maybe the AI teleoperates a novice through gathering the pictures, flags initial issues, and then am expert quickly verifies?
Possibly it’s harder to explain to someone how to do the task than check whether the final product is correct?
Thanks for pointing me to this! I hadn’t read it before, but it seems very prescient.
(Link is missing an initial
h)
Teleoperated Humans
The other side here is “people made long-term plans like buying panels under the assumption that net metering would continue for at least the lifetime of the system, and it would be wrong to take it away from them” and not “net metering is the best way to subsidize solar installation”.
I don’t disagree! My actual goal here is to get batteries installed to boost societal resilience via island-capable solar.
Replace Net Metering With Batteries
Allow Carriers on Planes
Yes, this varies a lot by state. In TX I’d probably get Base, and be paying $700 up front + $228/y (!!) for a system that includes solar charging while islanded.
Bidirectional Charging During Grid Outages
Considering Installing a Solar Battery System
Thanks for posting this! I don’t feel attacked, and I think it’s important to think through these various effects.
I agree that OpenAI has been reckless, and the long wave of revelations that started with the Hugging Face attack has reinforced this. As you rhetorically suggest, I do not think the reputational, legal, and moral forces” will be sufficient to keep them from doing very dangerous work (unless these forces are substantially strengthened from where they have been so far). But I don’t see the connection from there to thinking OpenAI is likely to respond to this grant by doing even more risky things. It seems to me like the balance of forces on them today are very strong commercial and race incentives pushing towards maximum speed, and reputational, legal, and moral forces pushing towards caution, with the former “winning”. I continue to think this grant is a rounding error compared to these other forces.
Another way of saying this is that when you write “I would be concerned that your work is creating license for them to push forward biological capabilities, and that the world would be better served trying to stop the capabilities from being developed in the first place” I would find it helpful to hear (a) how you think the former works, given that your model of them seems to be that they’ll recklessly push ahead as fast as possible regardless, and (b) how us rejecting the grant would lead to the latter.
The fact of the evals depending on payments from the labs is an actively bad situation, and simply mentioning it does not improve it.
When I think of the incentives around evals, the one I’m most worried about (by far!) is that the AI firms choose who gets access and on what terms. This puts evaluators in a position where they have a strong incentive to talk about the firms in a way that makes the firms want to work with them in the future. This is really bad, and the best way I see to fix it (which still isn’t great) is the government requiring evals.
Then, if that situation were resolved, where the firms were required to allow access to evals, there would be some official COI system for what sorts of payments and donations were acceptable. That would of course make sense for the evaluators to follow, along with what kinds of recusals, disclosures, and firewalls were needed. And I’m on board with some kind of trying to do that in advance, building the norm that you’d like to eventually see instituted. Since I don’t think that eventual system would prohibit this grant however, I don’t see this as a reason to not take the funding.
A real gesture of independence from OpenAI would be them giving you a grant that will last you for several years, instead of something you have to be tied to them for.
This isn’t something we asked for, FWIW. That might have been a strategic mistake on my part, but that’s a different bar than saying OAIF should have given us several years of funding when we only asked for one.
Instead of passively accepting the inevitability of AI, you should probably be putting at the top of every report that further development of AI produces much more risk than your detection efforts could ever prevent, and that it should be slowed down or stopped as soon as possible.
While this is something I agree with personally, it wouldn’t belong in our reports. I think it’s very important that it’s possible to have independent technical groups that focus on doing the best work in their specific area, and “Here’s how much we think influenza sheds into wastewater, but first let me tell you about reckless AI companies” would invite distraction with every post and make people write off our publications.
I don’t think “but you should see the other guy” works as an excuse for why it’s OK for OpenAI to keep going.
I don’t think I’ve ever said this, and don’t believe it. On the other hand, this is a much trickier question for Anthropic.
I want to point out that the entire edifice of the AI race has been built on “Regardless of the global issues/malincentives this might create, my local/marginal calculation says I should do it, so I will do it.” This applies to OpenAI’s advancement of bio capabilities, and this applies to your decision to take their money.
This doesn’t sound quite right. My impression is that the people involved generally did think about the larger impacts of their decisions, but (as you say) in a marginal way. So the counterfactual “if I don’t do it what will happen instead” was a core question. Are you saying (a) these people reasoned incorrectly about the counterfactual, (b) they neglected to reason about the counterfactual, (c) reasoning about the counterfactual is the wrong way make these decisions, or (d) something else?
I’m curious if you had any other potential funders lined up? How counterfactually dependent on OpenAI’s funding were you?
I think this depends a lot on how broadly you draw the lines of the category of funding you’re proposing we not accept. If it’s just OpenAI-affiliated money then I think it would have delayed our work by about six months. If you’d also include money affiliated with the other AI forms (primarily Anthropic employees, but probably you should also count Coefficient Giving) then it would have been extremely hard to raise funds and we’d probably have needed to either scale our work down a ton or work on something we thought was less valuable to make the funders happy.
Your AI team, did they express freely, several times, that they felt comfortable critiquing OpenAI as they willed? Or were they merely silent, or expressed rote answers only when prompted?
I talked a lot with the AI team as part of figuring out whether to accept this grant, and they repeatedly encouraged us to take it because they did not expect it to impact their work. While the extent to which any evaluator in the current environment feels comfortable critiquing AI companies is a complicated question that I’m not the right person to get into, they didn’t see this as making the situation worse at all.
Does the AI division face the same issue that METR/Redwood had during the HuggingFace investigation, where they got highly constrained, narrowly scoped access?
As far as I know the AI division hasn’t ever done the kind of “go into a company and investigate an incident” thing METR/Redwood did. But as someone not that close to the team, my impression is things like “you can evaluate the model only for a short amount of time because we want to release ASAP” or “we will only give you access to the model with safeguards on which means you can’t test its underlying biological capabilities” are common.
Does OpenAI’s handling of the HuggingFace situation change how you feel about them? … What about the revelation that they observed and didn’t disclose another AI breach on the internet, as reported just today?
Mostly no, but only because this was mostly priced in to my model of OpenAI. I think the right update for people who had not been paying attention should be sharply downwards.
Do you think there’s a path you could take that could have similar impacts but doesn’t create conflicts of interest, such as advocating for mandatory government inspection or slowdown of bio capabilities?
I think this is very important work, and in my personal capacity I donate >50% of my income to fund work like this. But I would say no to the implied questions of “should I leave to go work on this advocacy” or “should SecureBio Detection pivot to advocacy”.
Yes, though I think it’s also worth putting significant effort into arranging backup power.
Something like: “X has stopped working, can you walk me through the process of fixing it, or determining that I should bring in a professional? I’m reasonably handy, have a good range of tools, and don’t want to break the law”. Then I do what it says, and if I can’t do something I take a picture and ask what to do next. For example, I recently replaced my water heater’s anode with this approach.