METR has found that substantially different scaffolding is most effective for o-series models. I get the sense that they weren’t optimized for being effective multi-turn agents. At least, the o1 series wasn’t optimized for this, I think o3 may have been.
METR has found that substantially different scaffolding is most effective for o-series models. I get the sense that they weren’t optimized for being effective multi-turn agents. At least, the o1 series wasn’t optimized for this, I think o3 may have been.