On a different note, I’ve become increasingly worried that Resolution’s plan to automate alignment research will lead them to produce a lot of capabilities progress (since “automated alignment researcher” and “automated capabilities researcher” are such similar things to aim for).
You’ve been in favor of interpretability research in the past, wouldn’t that similarly increase capabilities?
You’ve been in favor of interpretability research in the past, wouldn’t that similarly increase capabilities?