Coherent use of language relies on tons of facts about (1) the world, and (2) the language. You’ll need to acquire this information, encode it in a human-understandable format, and provide it to the program. There’ll be too much information for one human to read, so it would need to be locally human-readable at multiple levels of abstraction. (So that you could parse the concept of an atom, a cell, a toe containing that cell, a human being, and a street full of human beings.) Without that sort of structure, it won’t really be comprehensible to humans.
I expect the structured encoding is the stumbling block. LLM training extracts and encodes world/language information, but encodes that information in massive matrices of inscrutable numbers.
In my experience building human-understandable models, successful models rely on high-level features with clear semantic meaning. Each feature should be understandable in isolation, and the model should only rely on limited interactions between those features. Otherwise, it becomes infeasible for a human to trace the logic and to reason about the model’s behavior. Language and vision rely on too many raw features with too much interaction, so you can’t directly build an understandable model on words or pixels. You can (and we did) build human-understandable models that sit on top of neural networks, letting the neural net parse messy raw inputs into a comprehensible feature for the human-understandable model to use. The result is a model that’s understandable unless you need to peek into the neural net’s workings, at which point NOPE.
For example, if you were trying to predict how much money a movie will make, you might want to know how hyped people seem to be about it. So you might run text analysis on articles/reviews/comments and convert that into a score (or a few scores) measuring the hype. A human-understandable model can’t directly use article text, but it can easily use a hype score as a feature.
Coherent use of language relies on tons of facts about (1) the world, and (2) the language. You’ll need to acquire this information, encode it in a human-understandable format, and provide it to the program. There’ll be too much information for one human to read, so it would need to be locally human-readable at multiple levels of abstraction. (So that you could parse the concept of an atom, a cell, a toe containing that cell, a human being, and a street full of human beings.) Without that sort of structure, it won’t really be comprehensible to humans.
I expect the structured encoding is the stumbling block. LLM training extracts and encodes world/language information, but encodes that information in massive matrices of inscrutable numbers.
In my experience building human-understandable models, successful models rely on high-level features with clear semantic meaning. Each feature should be understandable in isolation, and the model should only rely on limited interactions between those features. Otherwise, it becomes infeasible for a human to trace the logic and to reason about the model’s behavior. Language and vision rely on too many raw features with too much interaction, so you can’t directly build an understandable model on words or pixels. You can (and we did) build human-understandable models that sit on top of neural networks, letting the neural net parse messy raw inputs into a comprehensible feature for the human-understandable model to use. The result is a model that’s understandable unless you need to peek into the neural net’s workings, at which point NOPE.
For example, if you were trying to predict how much money a movie will make, you might want to know how hyped people seem to be about it. So you might run text analysis on articles/reviews/comments and convert that into a score (or a few scores) measuring the hype. A human-understandable model can’t directly use article text, but it can easily use a hype score as a feature.