If anyone sees this and is making a slop eval, let me know—it’d be nice to use it as a monitorability benchmark to see e.g. when a model oversells its work or half-asses the task, 1) is this visible to a monitor in the actions or CoT and 2) is this made plain and explicit to the user.
My guess is that 2) is a huge current monitorability failure.
If anyone sees this and is making a slop eval, let me know—it’d be nice to use it as a monitorability benchmark to see e.g. when a model oversells its work or half-asses the task, 1) is this visible to a monitor in the actions or CoT and 2) is this made plain and explicit to the user.
My guess is that 2) is a huge current monitorability failure.