There was a recent OpenAI announcement that they expect to spend 20% of the compute used for inference on on “monitoring compute” on that inference, at least in particularly important domains. This explicitly includes the inference that takes place in training.
This poses a natural questions: What percent of human training-time compute budget is used for monitoring humans for misbehavior?
Infants, children, teens, and adults spend a lot of time in “training”—at home, at school, at church, and so on. Some fraction of their “compute” learning is also spent by other people watching them and trying to track them for misbehavior. What’s the “monitoring compute” / “training compute” ratio?
Notably, we don’t want to just count moments when someone has caught someone misbehaving, but moments when someone is monitoring for “bad behavior.”
[I’m going to go with the awful / rough “fraction of total time spent monitoring” = “fraction of total personal compute spent monitoring”. On one hand, obviously your brain is doing a bunch of things other than monitoring while watching someone for misbehavior; but on the other hand, your brain is also spending some time figuring out how to watch for misbehavior when you’re not actively doing so, because it’s replaying shit through the hippocampus and consolidating memories. So… yeah.]
Relevant factors to consider:
(a) Parents: So my rough impression from my nieces and nephews is that my siblings spend on the order of… 1/50th to 1/200th of their total time actively correcting, thinking about how to correct, or considering how to correct their children. If we multiply by 6x for monitoring time we get 12% to 3% of their parent’s time spent on monitoring.
This matches up very approximately with mothers spending about ~100 minutes per day and fathers spending ~60 minutes per day monitoring children.
(Of course this investment decreases massively post childhood.)
(b) Teachers: Teachers spend time monitoring students for misbehavior. But overall high student / teacher ratios means this is probably comparatively negligible.
(c) Peers: Humans are strongly distinguished from other primates by the fact that they strongly tend to enforce norms on parties in a dispute, even when they are not a subject to the dispute. And of course, humans of all ages are monitoring those around them for unfairness / injustice vis-a-vis themselves. Overall, I want to say that maybe 2% − 20% of peers compute is spent doing something like monitoring their peers. This is probably also the most enduring compute, so far as it extends from childhood through time working.
Summing, and adding fudge factors, my guess is we get very approximately ~2%-30% of training-time compute spent on monitoring.
A further natural question is whether humans generally have a deep theory of how to stop misbehavior based on human nature. My guess is no.
And of course—whether we will need a deep theory of why AIs misbehave in order to stop them from misbehaving, or whether relatively simple measures following upon misbehavior-detection will work. People’s guesses differ here.
A very large fraction of white-collar work is based around the need for monitoring for misbehaviour of various sorts, even aspects that are not directly connected to that. A very large proportion is building and following processes that generate incentive structures to reduce the fraction of people misbehaving, even if the time spent following those processes is not in itself directly “monitoring” anyone in particular.
If we are talking about human to human interaction, I think we cannot speak of a binary divide between monitoring and not-monitoring. In some sense, humans are always passively monitoring. For example, if parents hear screaming from their children, they will often investigate/act even though they weren’t actively monitoring. More generally, monitoring and steering happens in all interactions slightly
There was a recent OpenAI announcement that they expect to spend 20% of the compute used for inference on on “monitoring compute” on that inference, at least in particularly important domains. This explicitly includes the inference that takes place in training.
This poses a natural questions: What percent of human training-time compute budget is used for monitoring humans for misbehavior?
Infants, children, teens, and adults spend a lot of time in “training”—at home, at school, at church, and so on. Some fraction of their “compute” learning is also spent by other people watching them and trying to track them for misbehavior. What’s the “monitoring compute” / “training compute” ratio?
Notably, we don’t want to just count moments when someone has caught someone misbehaving, but moments when someone is monitoring for “bad behavior.”
[I’m going to go with the awful / rough “fraction of total time spent monitoring” = “fraction of total personal compute spent monitoring”. On one hand, obviously your brain is doing a bunch of things other than monitoring while watching someone for misbehavior; but on the other hand, your brain is also spending some time figuring out how to watch for misbehavior when you’re not actively doing so, because it’s replaying shit through the hippocampus and consolidating memories. So… yeah.]
Relevant factors to consider:
(a) Parents: So my rough impression from my nieces and nephews is that my siblings spend on the order of… 1/50th to 1/200th of their total time actively correcting, thinking about how to correct, or considering how to correct their children. If we multiply by 6x for monitoring time we get 12% to 3% of their parent’s time spent on monitoring.
This matches up very approximately with mothers spending about ~100 minutes per day and fathers spending ~60 minutes per day monitoring children.
(Of course this investment decreases massively post childhood.)
(b) Teachers: Teachers spend time monitoring students for misbehavior. But overall high student / teacher ratios means this is probably comparatively negligible.
(c) Peers: Humans are strongly distinguished from other primates by the fact that they strongly tend to enforce norms on parties in a dispute, even when they are not a subject to the dispute. And of course, humans of all ages are monitoring those around them for unfairness / injustice vis-a-vis themselves. Overall, I want to say that maybe 2% − 20% of peers compute is spent doing something like monitoring their peers. This is probably also the most enduring compute, so far as it extends from childhood through time working.
Summing, and adding fudge factors, my guess is we get very approximately ~2%-30% of training-time compute spent on monitoring.
A further natural question is whether humans generally have a deep theory of how to stop misbehavior based on human nature. My guess is no.
And of course—whether we will need a deep theory of why AIs misbehave in order to stop them from misbehaving, or whether relatively simple measures following upon misbehavior-detection will work. People’s guesses differ here.
A very large fraction of white-collar work is based around the need for monitoring for misbehaviour of various sorts, even aspects that are not directly connected to that. A very large proportion is building and following processes that generate incentive structures to reduce the fraction of people misbehaving, even if the time spent following those processes is not in itself directly “monitoring” anyone in particular.
If we are talking about human to human interaction, I think we cannot speak of a binary divide between monitoring and not-monitoring. In some sense, humans are always passively monitoring. For example, if parents hear screaming from their children, they will often investigate/act even though they weren’t actively monitoring. More generally, monitoring and steering happens in all interactions slightly