We empirically measured this at the benchmark / research area level in Safetywashing (post here). About half the safety benchmarks we tested were highly correlated with upstream general capabilities, including “human preference alignment”/RLHF alignment benchmarks. We also show substantial confusion, where well-known “safety” goals and benchmarks are blurred, confused with, or used to advance capabilities.
Since many intuitive arguments (e.g. “alignment theory”) were not very productive and poorly predictive of empirical phenomena, we recommended safety benchmarks/areas report their capabilities correlation instead of arguing for their relevance verbally.
Much of our commentary on safetywashing mirrors this post’s observation. For example:
1) We comment on common flaws behind intuitive argumentation, including “safety through capabilities”:
In alignment theory, there is a tendency to theorize about what would be instrumentally useful for safety without adequately considering the need to improve the balance of safety and capabilities. This can lead to the promotion of capabilities research that happens to improve some safety benchmark scores (“safety via capabilities”), but in reality do not reduce overall risk.
2) In the paper, we also observe that safety research priorities are often dictated by {ad-hoc reasoning + popularity contests + AI corporations} rather than science through empirical measurement:
Research released by safety teams or famous safety researchers is often labeled as safety-relevant by default. Even if the work is one reframing and a new author list away from being perceived as a standard capabilities paper, the work is often “godfathered” in as a safety paper. The determination of whether an area is safety-relevant is often sociological rather than scientific.
We need a systematic, scientific way of identifying the research areas which will differentially contribute to safety, which we attempt to do in this paper. Researchers possibly need to converge on the right guidelines quickly. Without clear intellectual standards, random social accidents (e.g., what a random popular person supports, what a random grantmaker is “excited about,” etc.) will determine priorities.
We empirically measured this at the benchmark / research area level in Safetywashing (post here). About half the safety benchmarks we tested were highly correlated with upstream general capabilities, including “human preference alignment”/RLHF alignment benchmarks. We also show substantial confusion, where well-known “safety” goals and benchmarks are blurred, confused with, or used to advance capabilities.
Since many intuitive arguments (e.g. “alignment theory”) were not very productive and poorly predictive of empirical phenomena, we recommended safety benchmarks/areas report their capabilities correlation instead of arguing for their relevance verbally.
Much of our commentary on safetywashing mirrors this post’s observation. For example:
1) We comment on common flaws behind intuitive argumentation, including “safety through capabilities”:
2) In the paper, we also observe that safety research priorities are often dictated by {ad-hoc reasoning + popularity contests + AI corporations} rather than science through empirical measurement:
this makes a lot of sense. do you have any empirical data for 2025 and 2026 models and scales?