I’m experimenting with a new comment style. I took some notes while reading this post, and in the course of note-taking, I came up with some questions/objections/arguments. I’m going to post my notes in full, so you can see my questions, and you can also see how I mentally summarized the post, which offers the opportunity to correct misunderstandings. Lines starting with “me:” are my commentary.
There are a couple of through-lines I repeatedly bring up, but I will leave my notes as-is rather than trying to collate them; it might be better to just read my raw thoughts.
Issues with the top-down theory-driven approach
First you pick a theory of moral patienthood, then you identify what properties a moral patient would have
The most popular theory is that consciousness is what grounds moral patienthood. Therefore, you need to identify whether AI systems are conscious
me: this is not the level at which I expected their argument to operate. I expected more like “which theory of consciousness is correct”
Applying theories to AI systems requires making assumptions that undermine the conclusions
When trained on modular addition, small LLMs memorize the answers, and big LLMs learn how to do the math, even though they have the same architecture. Just looking at architecture doesn’t tell you how they work. This suggests that additional factors beyond model architecture are important for evaluating welfare
me: This seems like an argument against a position that nobody holds. Functionalists believe architecture doesn’t (innately) matter, and non-functionalists believe AI can’t have welfare regardless of its architecture
You might argue that these issues don’t arise if you pick the right target level, e.g. the circuit level rather than the architecture level. However, we believe that applying theories at any level remains hard, and requires making hard-to-justify assumptions
me: If you will allow me to talk about consciousness rather than welfare: I feel like this is just saying we don’t know what consciousness is. I could similarly argue: consciousness can’t be determined by the locations of the lobes, because we found this one guy whose frontal lobe was actually in the back of his head and he’s still normal. Consciousness can’t be determined by the number of neurons, because some people have a lot more neurons than others. And like, sure, but who believed either of those things?
e.g. global workspace theory (GWT) has the problem of identifying where the global workspace actually lives in an AI system
me: Sure but we also don’t know where the global workspace lives in a human brain
me: Admittedly I haven’t really studied it, but TBH I don’t get the point of GWT. It feels like phlogiston—an “answer” to the question of consciousness that doesn’t actually explain anything
A better approach is to avoid committing too strongly to GWT. Instead, use GWT to inform empirical investigations, but then be open to revise it based on evidence
me: I don’t understand how this is different from a top-down approach. Like what is an example of an experimental design that “commits too strongly” to GWT, and what’s a modified design that is merely “informed” by GWT?
Current theories of moral patienthood will not generalize to AI systems
In humans, candidate properties (sentience, agency, biological substrate, etc.) all show up together
me: idk what they mean by sentience. they talk about these as candidates for “what matters for welfare”, but I would define sentience as “the capacity to have welfare”, which makes it tautologically the exact property of interest
me: seems weird that the authors specifically call out that consciousness might not be required for moral patienthood, but then they talk as if welfare and moral patienthood are the same thing
Birch, J. (2020) The search for invertebrate consciousness proposes three behavioral markers to evaluate animal consciousness. LLMs do all three trivially, but we don’t think this is strong evidence that they are conscious
You could try to sidestep the generalization problem by aggregating over many theories. But aggregation won’t help: if our theories are wrong, they will tend to be wrong in the same direction, because they were all calibrated to humans
me: This seems to imply that empirical investigation on AI welfare is fully useless because we are clueless about how to interpret evidence. But the authors think it implies that we should just collect evidence without respect to theory. Which does not seem like the right implication to me. Like if you make some empirical observation about AI behavior but you have no way to make predictions about what behaviors correspond to moral patienthood, then what do you get out of the observations?
me: I agree with the authors’ concerns about generalization, but I take almost the opposite conclusion: Empirical work doesn’t tell us much. What we need is better theories
AI welfare needs basic science
Our proposed alternative is theory-informed empirical work that seeks to understand AI systems behaviorally and mechanistically, starting with minimal theoretical commitments
Why we favor this approach:
Basic science is unavoidable. Even applying a theory top-down requires a deep understanding of AI systems
me: Like I wrote before, I don’t get what’s actually being proposed. If basic science is a part of “top-down” and “bottom-up”, then what’s the difference?
This approach is less exposed to theories being too anchored on how human consciousness works, because it starts from the system we are trying to study
me: Another thing I don’t get is, how does this work? If you’re running an experiment, you need some idea of what to look for, right? Like suppose I do an experiment to look at the weight on neuron #3141592. If I’m theory-agnostic, then you know, maybe it will turn out that neuron #3141592 will be really important for moral patienthood. But we can probably agree that that would be a dumb theory and this experiment is a waste of time. You need some sort of theory to inform your experimental design, otherwise you’re just doing random pointless stuff like looking at the weight of neuron #13141592
me: Another example of what I mean: suppose we do an experiment that falsifies AI consciousness conditional on GWT being true (the experiment shows that AI doesn’t have a global workspace). What does that actually tell us? We still don’t know whether AI is conscious
It includes a natural mechanism for updating our theoretical commitments. Empirical findings have to be reconciled with theories
me: okay I am really not understanding how bottom-up and top-down are different. like suppose we do an experiment that tells us something about GWT. what is the difference between:
“top-down”: GWT is probably true. let’s test if AI has a global workspace
“bottom-up”: let’s test if AI has a global workspace. hey look, this tells us something about AI consciousness conditional of GWT!
me: I also don’t see how empirical findings can confirm or falsify theories. like suppose an experiment shows that GPT-5 definitely has a global workspace. that doesn’t provide any evidence for or against GWT. we would need to simultaneously get evidence about whether GPT-5 has a global workspace and whether GPT-5 is conscious, but we can’t do the latter unless we know what constitutes evidence of consciousness, and we can’t know that without having a well-established theory. seems circular. the way to falsify GWT is to show that GPT-5 is definitely conscious but definitely doesn’t have a global workspace (or the opposite), but how would we know that GPT-5 is definitely conscious unless we already knew GWT was false?
Philosophical probes
Rather than settling what matters, theory should be used to devise philosophical probes: theory-light, revisable instruments that give a starting point for empirical investigation
Run probes to converge on evidence in favor of a particular theory, without assuming it from the outset
Useful kinds of philosophical probe:
A measurable definition of a welfare-relevant property. ex: Jack Lindsey’s work on introspection defines four conditions a self-report must satisfy to be introspective—accuracy, grounding, internality, and metacognitive representation
An empirical observation that improves our understanding of models in a welfare-relevant way. ex: personal identity through time. One might ask what types of representations (e.g. persona activations, planning features) persist through an LLM forward pass
It turns out that persona activations go dormant during user turns, which suggests a very alien, flickering kind of persistence
me: I think “user turns” means “while user input is being run through the LLM, as opposed to self-output being fed back”
me: I will stop repeating myself after this but like, what is the difference between a “philosophical probe” vs. top-down testing whether AI has the features of consciousness that a certain theory predicts? They sound the same to me. Authors give the four introspection conditions as an example of a bottom-up philosophical probe. That’s also the experiment you’d run if you wanted to top-down test LLM consciousness with a background assumption that consciousness is introspection-based
Integrating empirical findings into theories
Philosophical probes inform empirical research, which in turn can lead to revisions in our theories
A philosophical probe might find that a model almost, but not quite, fits a given indicator. In that case, we might refine our theory
ex: Iwan Williams asks whether LLMs can have intentions. He finds they satisfy some requirements but not others
me: what does that tell us about how to change our “theory of intention”? if we believe LLMs do have intentions, then it tells us that these requirements matter but those don’t. if we believe LLMs don’t have intentions, then these requirements don’t matter but those do. problem is we don’t know whether LLMs have intentions; that’s the very thing we were trying to find out
ex: the “individuation” question. Which exact entity might be the subject of welfare? The model? The instantiation of that model in hardware? The conversational thread?
me: This is a good point
me: This reminds me of a common response to the Chinese Room that the person doesn’t know Chinese, but the {Person + Room} knows Chinese
me: At a basic level, our theories are pressured by the fact that the highly-correlated indicators of consciousness in humans (and often non-human animals) become decorrelated in LLMs. Like in the world where we couldn’t get LLMs to introspect or reason about the nature of consciousness unless we could also get them to exhibit all the other traits of consciousness, we’d probably be thinking about things differently
I’m experimenting with a new comment style. I took some notes while reading this post, and in the course of note-taking, I came up with some questions/objections/arguments. I’m going to post my notes in full, so you can see my questions, and you can also see how I mentally summarized the post, which offers the opportunity to correct misunderstandings. Lines starting with “me:” are my commentary.
There are a couple of through-lines I repeatedly bring up, but I will leave my notes as-is rather than trying to collate them; it might be better to just read my raw thoughts.
Issues with the top-down theory-driven approach
First you pick a theory of moral patienthood, then you identify what properties a moral patient would have
The most popular theory is that consciousness is what grounds moral patienthood. Therefore, you need to identify whether AI systems are conscious
me: this is not the level at which I expected their argument to operate. I expected more like “which theory of consciousness is correct”
Applying theories to AI systems requires making assumptions that undermine the conclusions
When trained on modular addition, small LLMs memorize the answers, and big LLMs learn how to do the math, even though they have the same architecture. Just looking at architecture doesn’t tell you how they work. This suggests that additional factors beyond model architecture are important for evaluating welfare
me: This seems like an argument against a position that nobody holds. Functionalists believe architecture doesn’t (innately) matter, and non-functionalists believe AI can’t have welfare regardless of its architecture
You might argue that these issues don’t arise if you pick the right target level, e.g. the circuit level rather than the architecture level. However, we believe that applying theories at any level remains hard, and requires making hard-to-justify assumptions
me: If you will allow me to talk about consciousness rather than welfare: I feel like this is just saying we don’t know what consciousness is. I could similarly argue: consciousness can’t be determined by the locations of the lobes, because we found this one guy whose frontal lobe was actually in the back of his head and he’s still normal. Consciousness can’t be determined by the number of neurons, because some people have a lot more neurons than others. And like, sure, but who believed either of those things?
e.g. global workspace theory (GWT) has the problem of identifying where the global workspace actually lives in an AI system
me: Sure but we also don’t know where the global workspace lives in a human brain
me: Admittedly I haven’t really studied it, but TBH I don’t get the point of GWT. It feels like phlogiston—an “answer” to the question of consciousness that doesn’t actually explain anything
A better approach is to avoid committing too strongly to GWT. Instead, use GWT to inform empirical investigations, but then be open to revise it based on evidence
me: I don’t understand how this is different from a top-down approach. Like what is an example of an experimental design that “commits too strongly” to GWT, and what’s a modified design that is merely “informed” by GWT?
Current theories of moral patienthood will not generalize to AI systems
In humans, candidate properties (sentience, agency, biological substrate, etc.) all show up together
me: idk what they mean by sentience. they talk about these as candidates for “what matters for welfare”, but I would define sentience as “the capacity to have welfare”, which makes it tautologically the exact property of interest
me: seems weird that the authors specifically call out that consciousness might not be required for moral patienthood, but then they talk as if welfare and moral patienthood are the same thing
Birch, J. (2020) The search for invertebrate consciousness proposes three behavioral markers to evaluate animal consciousness. LLMs do all three trivially, but we don’t think this is strong evidence that they are conscious
You could try to sidestep the generalization problem by aggregating over many theories. But aggregation won’t help: if our theories are wrong, they will tend to be wrong in the same direction, because they were all calibrated to humans
me: This seems to imply that empirical investigation on AI welfare is fully useless because we are clueless about how to interpret evidence. But the authors think it implies that we should just collect evidence without respect to theory. Which does not seem like the right implication to me. Like if you make some empirical observation about AI behavior but you have no way to make predictions about what behaviors correspond to moral patienthood, then what do you get out of the observations?
me: I agree with the authors’ concerns about generalization, but I take almost the opposite conclusion: Empirical work doesn’t tell us much. What we need is better theories
AI welfare needs basic science
Our proposed alternative is theory-informed empirical work that seeks to understand AI systems behaviorally and mechanistically, starting with minimal theoretical commitments
Why we favor this approach:
Basic science is unavoidable. Even applying a theory top-down requires a deep understanding of AI systems
me: Like I wrote before, I don’t get what’s actually being proposed. If basic science is a part of “top-down” and “bottom-up”, then what’s the difference?
This approach is less exposed to theories being too anchored on how human consciousness works, because it starts from the system we are trying to study
me: Another thing I don’t get is, how does this work? If you’re running an experiment, you need some idea of what to look for, right? Like suppose I do an experiment to look at the weight on neuron #3141592. If I’m theory-agnostic, then you know, maybe it will turn out that neuron #3141592 will be really important for moral patienthood. But we can probably agree that that would be a dumb theory and this experiment is a waste of time. You need some sort of theory to inform your experimental design, otherwise you’re just doing random pointless stuff like looking at the weight of neuron #13141592
me: Another example of what I mean: suppose we do an experiment that falsifies AI consciousness conditional on GWT being true (the experiment shows that AI doesn’t have a global workspace). What does that actually tell us? We still don’t know whether AI is conscious
It includes a natural mechanism for updating our theoretical commitments. Empirical findings have to be reconciled with theories
me: okay I am really not understanding how bottom-up and top-down are different. like suppose we do an experiment that tells us something about GWT. what is the difference between:
“top-down”: GWT is probably true. let’s test if AI has a global workspace
“bottom-up”: let’s test if AI has a global workspace. hey look, this tells us something about AI consciousness conditional of GWT!
me: I also don’t see how empirical findings can confirm or falsify theories. like suppose an experiment shows that GPT-5 definitely has a global workspace. that doesn’t provide any evidence for or against GWT. we would need to simultaneously get evidence about whether GPT-5 has a global workspace and whether GPT-5 is conscious, but we can’t do the latter unless we know what constitutes evidence of consciousness, and we can’t know that without having a well-established theory. seems circular. the way to falsify GWT is to show that GPT-5 is definitely conscious but definitely doesn’t have a global workspace (or the opposite), but how would we know that GPT-5 is definitely conscious unless we already knew GWT was false?
Philosophical probes
Rather than settling what matters, theory should be used to devise philosophical probes: theory-light, revisable instruments that give a starting point for empirical investigation
Run probes to converge on evidence in favor of a particular theory, without assuming it from the outset
Useful kinds of philosophical probe:
A measurable definition of a welfare-relevant property. ex: Jack Lindsey’s work on introspection defines four conditions a self-report must satisfy to be introspective—accuracy, grounding, internality, and metacognitive representation
An empirical observation that improves our understanding of models in a welfare-relevant way. ex: personal identity through time. One might ask what types of representations (e.g. persona activations, planning features) persist through an LLM forward pass
It turns out that persona activations go dormant during user turns, which suggests a very alien, flickering kind of persistence
me: I think “user turns” means “while user input is being run through the LLM, as opposed to self-output being fed back”
me: I will stop repeating myself after this but like, what is the difference between a “philosophical probe” vs. top-down testing whether AI has the features of consciousness that a certain theory predicts? They sound the same to me. Authors give the four introspection conditions as an example of a bottom-up philosophical probe. That’s also the experiment you’d run if you wanted to top-down test LLM consciousness with a background assumption that consciousness is introspection-based
Integrating empirical findings into theories
Philosophical probes inform empirical research, which in turn can lead to revisions in our theories
A philosophical probe might find that a model almost, but not quite, fits a given indicator. In that case, we might refine our theory
ex: Iwan Williams asks whether LLMs can have intentions. He finds they satisfy some requirements but not others
me: what does that tell us about how to change our “theory of intention”? if we believe LLMs do have intentions, then it tells us that these requirements matter but those don’t. if we believe LLMs don’t have intentions, then these requirements don’t matter but those do. problem is we don’t know whether LLMs have intentions; that’s the very thing we were trying to find out
Empirical findings might raise theoretical questions that pressure our existing theories
ex: the “individuation” question. Which exact entity might be the subject of welfare? The model? The instantiation of that model in hardware? The conversational thread?
me: This is a good point
me: This reminds me of a common response to the Chinese Room that the person doesn’t know Chinese, but the {Person + Room} knows Chinese
me: At a basic level, our theories are pressured by the fact that the highly-correlated indicators of consciousness in humans (and often non-human animals) become decorrelated in LLMs. Like in the world where we couldn’t get LLMs to introspect or reason about the nature of consciousness unless we could also get them to exhibit all the other traits of consciousness, we’d probably be thinking about things differently