That question sounds confused and poses the issues as a binary yes/no question. A model can certainly engage in some reasoning on it’s own. That however does not tell you how good it’s evaluation skills are. They are probably not perfect but also not non-existent.
Would this model even be able to evaluate its alignment itself? (or is it just the result of scraping what other people think about ai alignment?)
That question sounds confused and poses the issues as a binary yes/no question. A model can certainly engage in some reasoning on it’s own. That however does not tell you how good it’s evaluation skills are. They are probably not perfect but also not non-existent.