IIRC GPT-5.5 has a chess rating of ~1600 Elo. Suppose that you asked it to write a story about two grandmasters of chess and to describe the entire game. Then the game would either be written by a chess bot or reveal GPT-5.5′s lack of the skill. If we replace chess with forecasting or writing the scenarios,[1] then you are likely to get shitty results like METR’s failed attempt to have an unrevealed model write out a report on autonomous replication.
Or doing philosophy, but AI-generated philosophy is harder to evaluate and could be confounded by the cultural hegemon’s priors, RLHF favoring sycophantic philosophy, etc.
Suppose that you hired an editor to write a paragraph in your post. Doesn’t this require you to disclose who actually wrote the paragraph?
IIRC GPT-5.5 has a chess rating of ~1600 Elo. Suppose that you asked it to write a story about two grandmasters of chess and to describe the entire game. Then the game would either be written by a chess bot or reveal GPT-5.5′s lack of the skill. If we replace chess with forecasting or writing the scenarios,[1] then you are likely to get shitty results like METR’s failed attempt to have an unrevealed model write out a report on autonomous replication.
Or doing philosophy, but AI-generated philosophy is harder to evaluate and could be confounded by the cultural hegemon’s priors, RLHF favoring sycophantic philosophy, etc.