The headline claim seems true for people invested in the rationalist community but I don’t understand the case you’re making. It seems like the argument here is that the lineage of Robin Hanson->yourself->LW in general is upstream of people in general treating this phrase as a reason to cultivate the corresponding mental habit? As in:
Robin Hanson is the person I know who went around saying to people that they ought not to update in a predictable direction, as a consequence of the martingale principle underlying EMH, applied to their personal epistemic lives.
This is probably true for people alive today, because communities using probability as a formal framework for refining their epistemics are rare, but I’m not convinced that this is relevant to the headline claim. As a child I remember realizing this within an hour of being taught Bayes’ Theorem, and while it’s possible that my math camp instructors were themselves influenced by yourself or Hanson and guided me in this direction subtly, I doubt it. I was also not particularly precocious among precocious children, such that I doubt my experience here is rare. That is to say, I think you may be overestimating how basic this fact is, and I think the reason it’s hard to find it formally written up in the relevant contexts is that it’s so obvious as to be embarrassing to try to explain to anybody who already has relevant skill in probability. Many people have engaged with probability as a formal discipline in history, very few of them have taken its simplest exercises as seriously as you have in their writing, and the discipline is old enough that those who did were mostly not writing in such a well-preserved medium. (This is a judgment but I’m unsure of its valence. Taking simple things seriously is often a good idea and often naval-gazing and I struggle to differentiate the two.)
By comparison, not many people have thought about AI safety as a formal discipline, and so I think analogous claims in that field are much more likely to be true—the relevant ideas are legitimately understudied, very few capable people have tried to build useful frameworks for that project, and you’re correct to claim a bottleneck position in that story. Indeed, very few people took AI safety seriously until a couple of years ago, and so ~all of the expertise is held in a highly-noncomformist social group with a ~unique and controversial culture around sex, emotional well-being, communication, values concerning future people and the futures of current people, etc. This is in my view the main problem for AI safety communicators right now! If this bottleneck didn’t exist, neither would this problem. No comparable thing seems true for probability.
[Of course, if the target audience for this post is “people who were first introduced to Bayesian epistemics from the writings of Eliezer Yudkowsky” rather than “people who might read a front-page post on LW” then I’m entirely incorrect and this comment can safely be deleted.]
As a child I remember realizing this within an hour of being taught Bayes’ Theorem
In contrast, it took me about 10 hours of reading and rereading, carefully parsing, and doing exercises, when I was about 28, to get to the point of properly understanding Bayes’ Theorem. (This is after having read the sequences years earlier, and having been embedded in the Bay area rationalist community, and working at CFAR for half a decade.)
I do not think I would have hit on this insight independently, without it being pointed out to me, or without my doing a bunch of independent thinking about technical epistemology that I would have been very unlikely to do, without someone like Eliezer pointing me in the right direction.
That is to say, I think you may be overestimating how basic this fact is
I think you might be underestimating how non-intuitive even the basic facts are to almost everyone (possibly even to people you would generally regard as a “smart person”, though that is indeed less obvious).
Since the Sequences expanded probabilistic thinking to a broader audience, who were mostly not reading the works of the probabilistic thinkers of Eld or even thinking about the math by themselves much, the Sequences would be the source of probabilistic principles among almost all of that community, even if they were capable of coming up with conservation of expected evidence within a couple hours of thinking about basic probability.
I also find it intriguing that speck did not say “It took me a few minutes to figure out it out after learning about conditional probability as a kid?”, which presumably they learned before Bayes law, and is the plausible pathway through which a significant chunk of the new probabilistic thinkers could’ve rediscovered it.
For myself, I think I was familiar with how conditional probability worked before I knew algebra, at a level sufficient to solve a simple problem embodying P(X) = P(E)P(X|E) + P(-E)P(X|-E), and people who did math competitions as kids are liable to passively have that knowledge—yet at no point did I realize that the expected evidence is conserved, even after I learned enough algebra to write that equation down. I would give myself a 10% chance of having come up with it had I had the concept of updating beliefs in response to evidence. As a benchmark, Bayes Law and Conservation of Expected Evidence were easy for me to learn from the Sequences at ~13 - that is, someone with the math skill to find them easy as a kid is still probably not going to come up with CoEE!
I also find it intriguing that speck did not say “It took me a few minutes to figure out it out after learning about conditional probability as a kid?”, which presumably they learned before Bayes law, and is the plausible pathway through which a significant chunk of the new probabilistic thinkers could’ve rediscovered it.
I recall learning these simultaneously, but it’s possible I had seen and not internalized conditional probability before. The pathway as I remember it was: broad description of what it means for an event to have probability p, description of what it means for an event to have probability p conditional on E, probabilities of Boolean combinations of events, derivation of Bayes, and then in exercises derivation of and application of the Law of Total Probability. After you’ve used P(X) = P(E)P(X|E) + P(-E)P(X|-E) like twice, I think there’s a reasonable chance of noticing that you keep computing 0.6P(X|E) + 0.4P(X|~E) or 0.2P(X|E) + 0.8P(X|~E), thinking a little bit about units, recognizing this as a weighted average, and then there’s nothing really to observe.
If people are learning about probability in ways other than writing down lots of toy numbers and doing arithmetic, or aren’t making the leap of treating your own credences as probabilities and pieces of evidences as events to condition on, then it’s a bit more plausible to me that figuring this out is difficult. For example, @Eli Tyre, you write:
In contrast, it took me about 10 hours of reading and rereading, carefully parsing, and doing exercises, when I was about 28, to get to the point of properly understanding Bayes’ Theorem. (This is after having read the sequences years earlier, and having been embedded in the Bay area rationalist community, and working at CFAR for half a decade.)
and depending on what you mean by ‘exercises’ here, one possible explanation for the difference in our experiences here is that reading, parsing, participating in the rationalist community, reading the sequences, and working at CFAR are all something other than mechanically computing dozens of toy probabilities until the steps become muscle memory. Another possible explanation is that I’m just having a theory of mind failure, which is the main thing I’m trying to settle here.
It’s also plausible that by ‘Bayesians don’t predictably update’ you mean ‘conservation of expected evidence’ you mean ‘the law of total probability’? My general model here is that the law of total probability is a fairly obvious fact, the sort you couldn’t hope to miss by doing enough rote computation on the relevant problems. Meanwhile ‘conservation of expected evidence’ is a way of interpreting that law by treating your credence in an event as a random variable distributed against the evidence you might see, and ‘Bayesians don’t predictably update’ is the mental habit of ensuring that your beliefs represent truth values by ensuring that they follow this law. Is this roughly in line with the way y’all are using these terms?
I meant to use “conservation of expected evidence” as that interpretation of the law of total probability, and was using no particular word or phrase for the mental habit (which I would’ve just called “the mental habit of trying to do this”).
Personally I’m unsure whether I’d have noticed it if you erased it from my mind but kept my math ability and had me toy around with probabilities, simply because most of the time I’m using the law of total probability I am not in fact interpreting the conditioned variable as evidence for a hypothesis. I think the best shot I have at thinking to look in that direction is if I got into an argument with someone and then ended up reinventing the idea that you shouldn’t expect to update on average via economics reasoning on stock prices.
...I wonder how many more people would know Bayes law if someone just convinced the common national math competition orgs to put questions using it on their tests… after all, that’s how I ended up passively having an intuitive understanding of expected values and conditional probability.
the reason it’s hard to find it formally written up in the relevant contexts is that it’s so obvious as to be embarrassing to try to explain to anybody who already has relevant skill in probability
I disagree. Up until the proof-based courses of undergraduate upperclassmen, math education involves an enormous amount of hand-holding. This is not a boast—I’m not saying that the classes are easy. What I mean is that students aren’t expected to reach any conclusions on their own, and every conclusion that they’re presented with is reinforced through practice exercises and exam questions, often including word problems.
This depends on the education tract, but broadly this is true and doesn’t seem relevant to what I’m saying here unless I’m missing something. Most people don’t ever develop any relevant skill in probability! Most people don’t treat math as an epistemic foundation, and so wouldn’t think the headline claim is obvious or meaningful. That is to say: I agree that math education involves a lot of handholding, even at an upperclassman level, but that’s because math education is targeted toward people who don’t really think about math outside of the classroom. People who do think about math outside of the classroom, for example who think of their beliefs as probabilities, are much more likely to eventually find their way to these conclusions.
The headline claim seems true for people invested in the rationalist community but I don’t understand the case you’re making. It seems like the argument here is that the lineage of Robin Hanson->yourself->LW in general is upstream of people in general treating this phrase as a reason to cultivate the corresponding mental habit? As in:
This is probably true for people alive today, because communities using probability as a formal framework for refining their epistemics are rare, but I’m not convinced that this is relevant to the headline claim. As a child I remember realizing this within an hour of being taught Bayes’ Theorem, and while it’s possible that my math camp instructors were themselves influenced by yourself or Hanson and guided me in this direction subtly, I doubt it. I was also not particularly precocious among precocious children, such that I doubt my experience here is rare. That is to say, I think you may be overestimating how basic this fact is, and I think the reason it’s hard to find it formally written up in the relevant contexts is that it’s so obvious as to be embarrassing to try to explain to anybody who already has relevant skill in probability. Many people have engaged with probability as a formal discipline in history, very few of them have taken its simplest exercises as seriously as you have in their writing, and the discipline is old enough that those who did were mostly not writing in such a well-preserved medium. (This is a judgment but I’m unsure of its valence. Taking simple things seriously is often a good idea and often naval-gazing and I struggle to differentiate the two.)
By comparison, not many people have thought about AI safety as a formal discipline, and so I think analogous claims in that field are much more likely to be true—the relevant ideas are legitimately understudied, very few capable people have tried to build useful frameworks for that project, and you’re correct to claim a bottleneck position in that story. Indeed, very few people took AI safety seriously until a couple of years ago, and so ~all of the expertise is held in a highly-noncomformist social group with a ~unique and controversial culture around sex, emotional well-being, communication, values concerning future people and the futures of current people, etc. This is in my view the main problem for AI safety communicators right now! If this bottleneck didn’t exist, neither would this problem. No comparable thing seems true for probability.
[Of course, if the target audience for this post is “people who were first introduced to Bayesian epistemics from the writings of Eliezer Yudkowsky” rather than “people who might read a front-page post on LW” then I’m entirely incorrect and this comment can safely be deleted.]
In contrast, it took me about 10 hours of reading and rereading, carefully parsing, and doing exercises, when I was about 28, to get to the point of properly understanding Bayes’ Theorem. (This is after having read the sequences years earlier, and having been embedded in the Bay area rationalist community, and working at CFAR for half a decade.)
I do not think I would have hit on this insight independently, without it being pointed out to me, or without my doing a bunch of independent thinking about technical epistemology that I would have been very unlikely to do, without someone like Eliezer pointing me in the right direction.
I think you might be underestimating how non-intuitive even the basic facts are to almost everyone (possibly even to people you would generally regard as a “smart person”, though that is indeed less obvious).
Since the Sequences expanded probabilistic thinking to a broader audience, who were mostly not reading the works of the probabilistic thinkers of Eld or even thinking about the math by themselves much, the Sequences would be the source of probabilistic principles among almost all of that community, even if they were capable of coming up with conservation of expected evidence within a couple hours of thinking about basic probability.
I also find it intriguing that speck did not say “It took me a few minutes to figure out it out after learning about conditional probability as a kid?”, which presumably they learned before Bayes law, and is the plausible pathway through which a significant chunk of the new probabilistic thinkers could’ve rediscovered it.
For myself, I think I was familiar with how conditional probability worked before I knew algebra, at a level sufficient to solve a simple problem embodying P(X) = P(E)P(X|E) + P(-E)P(X|-E), and people who did math competitions as kids are liable to passively have that knowledge—yet at no point did I realize that the expected evidence is conserved, even after I learned enough algebra to write that equation down. I would give myself a 10% chance of having come up with it had I had the concept of updating beliefs in response to evidence. As a benchmark, Bayes Law and Conservation of Expected Evidence were easy for me to learn from the Sequences at ~13 - that is, someone with the math skill to find them easy as a kid is still probably not going to come up with CoEE!
I recall learning these simultaneously, but it’s possible I had seen and not internalized conditional probability before. The pathway as I remember it was: broad description of what it means for an event to have probability p, description of what it means for an event to have probability p conditional on E, probabilities of Boolean combinations of events, derivation of Bayes, and then in exercises derivation of and application of the Law of Total Probability. After you’ve used P(X) = P(E)P(X|E) + P(-E)P(X|-E) like twice, I think there’s a reasonable chance of noticing that you keep computing 0.6P(X|E) + 0.4P(X|~E) or 0.2P(X|E) + 0.8P(X|~E), thinking a little bit about units, recognizing this as a weighted average, and then there’s nothing really to observe.
If people are learning about probability in ways other than writing down lots of toy numbers and doing arithmetic, or aren’t making the leap of treating your own credences as probabilities and pieces of evidences as events to condition on, then it’s a bit more plausible to me that figuring this out is difficult. For example, @Eli Tyre, you write:
and depending on what you mean by ‘exercises’ here, one possible explanation for the difference in our experiences here is that reading, parsing, participating in the rationalist community, reading the sequences, and working at CFAR are all something other than mechanically computing dozens of toy probabilities until the steps become muscle memory. Another possible explanation is that I’m just having a theory of mind failure, which is the main thing I’m trying to settle here.
It’s also plausible that by ‘Bayesians don’t predictably update’ you mean ‘conservation of expected evidence’ you mean ‘the law of total probability’? My general model here is that the law of total probability is a fairly obvious fact, the sort you couldn’t hope to miss by doing enough rote computation on the relevant problems. Meanwhile ‘conservation of expected evidence’ is a way of interpreting that law by treating your credence in an event as a random variable distributed against the evidence you might see, and ‘Bayesians don’t predictably update’ is the mental habit of ensuring that your beliefs represent truth values by ensuring that they follow this law. Is this roughly in line with the way y’all are using these terms?
I meant to use “conservation of expected evidence” as that interpretation of the law of total probability, and was using no particular word or phrase for the mental habit (which I would’ve just called “the mental habit of trying to do this”).
Personally I’m unsure whether I’d have noticed it if you erased it from my mind but kept my math ability and had me toy around with probabilities, simply because most of the time I’m using the law of total probability I am not in fact interpreting the conditioned variable as evidence for a hypothesis. I think the best shot I have at thinking to look in that direction is if I got into an argument with someone and then ended up reinventing the idea that you shouldn’t expect to update on average via economics reasoning on stock prices.
...I wonder how many more people would know Bayes law if someone just convinced the common national math competition orgs to put questions using it on their tests… after all, that’s how I ended up passively having an intuitive understanding of expected values and conditional probability.
I disagree. Up until the proof-based courses of undergraduate upperclassmen, math education involves an enormous amount of hand-holding. This is not a boast—I’m not saying that the classes are easy. What I mean is that students aren’t expected to reach any conclusions on their own, and every conclusion that they’re presented with is reinforced through practice exercises and exam questions, often including word problems.
This depends on the education tract, but broadly this is true and doesn’t seem relevant to what I’m saying here unless I’m missing something. Most people don’t ever develop any relevant skill in probability! Most people don’t treat math as an epistemic foundation, and so wouldn’t think the headline claim is obvious or meaningful. That is to say: I agree that math education involves a lot of handholding, even at an upperclassman level, but that’s because math education is targeted toward people who don’t really think about math outside of the classroom. People who do think about math outside of the classroom, for example who think of their beliefs as probabilities, are much more likely to eventually find their way to these conclusions.