Some notes on what sloppy AIs might look like in the next few years:
Maybe they really want apparent success or score, or maybe they have a mess of shallower heuristics and drives that result in slop. Similar to how current models write prose that’s (by pre-LLM standards) really high-quality in all of the superficial, easy-to-check ways like having easy-to-read flow and good turns-of-phrase, but whose substance and longer-run coherence etc is subtly bad. This particular problem will probably get better, but it might just push the “slop horizons” out, so it would take us more time to notice how the AI’s output is flawed.
E.g., it chooses the wrong benchmarks for AI R&D. For AI capabilities, we probably notice these issues in a few weeks or months when the deployed model has some measurable deficits. But in AI safety we don’t get this feedback as much (since the AIs might not be capable enough to be dangerously misaligned yet, or we might never notice if the models are misaligned), and picking benchmarks for AI safety is a harder problem because there’s a bigger distributional gap with bigger external validity concerns.
We’ll probably be pretty aware of this problem at the time but might struggle to elicit better labor for solving alignment.
Some notes on what sloppy AIs might look like in the next few years:
Maybe they really want apparent success or score, or maybe they have a mess of shallower heuristics and drives that result in slop. Similar to how current models write prose that’s (by pre-LLM standards) really high-quality in all of the superficial, easy-to-check ways like having easy-to-read flow and good turns-of-phrase, but whose substance and longer-run coherence etc is subtly bad. This particular problem will probably get better, but it might just push the “slop horizons” out, so it would take us more time to notice how the AI’s output is flawed.
E.g., it chooses the wrong benchmarks for AI R&D. For AI capabilities, we probably notice these issues in a few weeks or months when the deployed model has some measurable deficits. But in AI safety we don’t get this feedback as much (since the AIs might not be capable enough to be dangerously misaligned yet, or we might never notice if the models are misaligned), and picking benchmarks for AI safety is a harder problem because there’s a bigger distributional gap with bigger external validity concerns.
We’ll probably be pretty aware of this problem at the time but might struggle to elicit better labor for solving alignment.