I’d love to also test this against open source models, or more broadly have a benchmark of how well different models understand alignment.
I’d love to also test this against open source models, or more broadly have a benchmark of how well different models understand alignment.