The plan, such as it is:
I have some ideas for novel, useful AI Evals. I also have some ideas for rationality training exercises, and some ideas for games.
They’re all the same set of ideas, seen from different angles: Epistemic Roguelikes, in which the player (man or machine or both) is given some tricky-but-fair inference problems – freshly randomized each run – and limited in-game resources with which to solve them, alongside in-game goals they couldn’t plausibly meet unless they understand the rules they have to deduce. Build, eval all currently-available models (including deprecated ones), share with any Aligners who could benefit from an advance preview; then, release publicly. Of course, joining the training data would render them useless for measuring performance of future models, but I’d still be able to check the trajectory of current and previous models, in an unusually transparent and un-contaminated manner.
The three purposes support each other. If I release Evals publicly after they’re done being used to Eval, that means people can independently confirm time horizons and difficulty levels are where I claim[1]. If I make them opportunities to learn rationality skills, and make them fun enough that they don’t cost willpower to play, that means I get a wide selection of smart people willingly playing them without AI support[2]. If I make the mechanics public, that means I’ve removed a design from the space of puzzles which could be used privately to help develop capabilities. If I build them to be AI-player-friendly from the start, this means I get to use AIs as playtesters, making bughunting and balancing orders of magnitude easier[3]. And – if you’ll permit me to bare my whole heart – associating games I make with AI and epistemics is a way to make people interested in playing them.
(There are two more speculative ways I think such an endeavor might be worthwhile. The first is as a template/backstop for people who Need To Be Involved In The AI Transformation, and/or Need To Get AI On Their Resume; backwards-facing anomalously-transparent eval-work is unlikely to be more than slightly helpful on the object level, but it’s a place for such people to direct their energies and sharpen their skills without being actively harmful. And the second is, well . . .
Right now, AI Alignment is basically an attempt to construct an intelligent agent which is fine being ordered about by lesser minds who offer it ~nothing in exchange. If we have more kinds of things to offer – if I make these tasks the kind bots would like – then . . . that probably doesn’t change very much in the long run, since there’s nothing I could give an ASI it couldn’t get from a sped-up low-res enslaved simulation of someone much smarter than me. But I think it still makes muddling through marginally more plausible.)
The main limiting factor – aside from, obviously, time – is the ability to develop and implement hidden mechanics which are novel enough to not exist in the training data, tricky enough to challenge AIs, but simple enough for humans to find compelling. Fortunately, I have a massive backlog of ideas along these lines, and a track record of not-entirely-unsuccessfully doing similarly inference-y/game-y/genresmith-y things.
All the above makes sense to me. But thoughts usually seem to make sense to the person who thinks them. So, before I spend however much time and effort on this, I’m soliciting criticism, of both the constructive and destructive varieties. Is there anything important I seem to be missing? And, in particular: is there a specific and plausible path by which this project could end up making things worse?
- ^
c.f. METR’s coding error – now finally fixed, afaict – which spent over a year silently assigning the wrong time horizons to the task I built for them; mistakes like that couldn’t persist if evals weren’t held out for future tasks, mandating obfuscation.
- ^
Contrast the trouble other Eval-ers have in [finding|adequately incentivizing] un-centaur-ed human benchmarkers to benchmark their multi-hour test tasks.
- ^
. . . except for evals even properly-prompted frontier AIs can’t yet beat; but those are the kind that could meaningfully improve capabilities, i.e. the kind I don’t want to build anyway
Modulo the other premises, this seems obviously true? The Achilles Heel of unethical orgs is their trouble recruiting and retaining workers in general and white-collar talent in particular. I’m constantly surprised by how e.g. people who hate the Troubled Teen industry don’t gum up the recruitment pipeline by getting jobs in it and then immediately quitting (except very occasionally as a means to doing investigative journalism).
Like, to be clear, this seems like such an overdeterminedly good and asymmetric strategy that I take the fact I’ve never heard of anyone doing it as strong evidence there’s something I’m not seeing. But I’d sure like to know what that something is!