If “realistically simulating human impacts in silico” means something like high-fidelity models of how humans respond psychologically to AI interaction, that seems like one of the higher-risk cases of safety research also being capabilities research. A good sim for “how AI affects humans psychologically” is a good sim for “how to affect humans psychologically, using AI”.
The hardest part of CogSec evals is figuring out an eval meaningful enough to detect dangerous effects, while not being meaningful enough to hill climb on. Maybe eval is too benchmark coded, when actually what we want is comprehensive user surveying, similar to Anthropic’s What 81,000 people want from AI, but conducted independently, tied to usage data in a privacy-respecting way. One of the genuinely good uses of AI here may be qualitative research at scale: preserving the nuance of how people feel about and relate to AI, rather than immediately collapsing it into 1–5 survey responses.
I agree that this is an important externality and it’s something I think about a fair amount.
My current view on this is:
We can roughly decompose into two questions:
(A) “does AI behavior X have psychological effect Y on humans” and
(B) “how much of a propensity does this AI system have to exhibit behavior X?”
We will typically answer (A) with a combination of existing psychology literature, longitudinal studies, and intuition from domain experts.
We will typically answer (B) with in silico simulations
We will also use longitudinal studies to sanity check that the answers to (A) and (B) actually compose as expected.
To answer (B), you need simulations that are similar enough to humans to elicit similar behaviors from the language model. But these simulations are short-term, not long-term, so they don’t need to simulate the long-term effects on humans.
An unscrupulous company could potentially use the simulations from (B) to optimize for behaviors that elicit the desired short-term responses from humans. But since we’re looking at short-term effects, they could have already optimized directly on their pool of users; there isn’t much uplift from a simulation.
The main advantage of simulations is that they (1) give you apples-to-apples comparisons across different models, and (2) let you make measurements even if you don’t have a ton of user traffic to draw on. Both of these differentially help evaluators compared to large companies.
If “realistically simulating human impacts in silico” means something like high-fidelity models of how humans respond psychologically to AI interaction, that seems like one of the higher-risk cases of safety research also being capabilities research. A good sim for “how AI affects humans psychologically” is a good sim for “how to affect humans psychologically, using AI”.
The hardest part of CogSec evals is figuring out an eval meaningful enough to detect dangerous effects, while not being meaningful enough to hill climb on. Maybe eval is too benchmark coded, when actually what we want is comprehensive user surveying, similar to Anthropic’s What 81,000 people want from AI, but conducted independently, tied to usage data in a privacy-respecting way. One of the genuinely good uses of AI here may be qualitative research at scale: preserving the nuance of how people feel about and relate to AI, rather than immediately collapsing it into 1–5 survey responses.
I agree that this is an important externality and it’s something I think about a fair amount.
My current view on this is:
We can roughly decompose into two questions: (A) “does AI behavior X have psychological effect Y on humans” and (B) “how much of a propensity does this AI system have to exhibit behavior X?”
We will typically answer (A) with a combination of existing psychology literature, longitudinal studies, and intuition from domain experts.
We will typically answer (B) with in silico simulations
We will also use longitudinal studies to sanity check that the answers to (A) and (B) actually compose as expected.
To answer (B), you need simulations that are similar enough to humans to elicit similar behaviors from the language model. But these simulations are short-term, not long-term, so they don’t need to simulate the long-term effects on humans.
An unscrupulous company could potentially use the simulations from (B) to optimize for behaviors that elicit the desired short-term responses from humans. But since we’re looking at short-term effects, they could have already optimized directly on their pool of users; there isn’t much uplift from a simulation.
The main advantage of simulations is that they (1) give you apples-to-apples comparisons across different models, and (2) let you make measurements even if you don’t have a ton of user traffic to draw on. Both of these differentially help evaluators compared to large companies.