Yes, I certainly agree that our compressor is currently more universally applicable, which is exactly why I am excited about seeing this sort of auto-compressing behavior in models like Mythos (and [now Sonnet 5 also](https://thezvi.substack.com/i/204364347/illegible-thinking-645)).
To me, and keep in mind I’m not a mathematician, the distribution that mathematical abstractions compress in this framing has always been a bit mysterious.
The basic idea, as I think of it, is this: in mathematics we have objects and we have facts about those objects. For example, if you consider the integers, , we have a pretty good sense of this object. There is an addition and a multiplication on , objects are invertible under addition (meaning for all there exists such that ), but not under multiplication. One can continue to list out facts. Now, is also a group under addition. Because is a group, it inherits all the facts that we know about groups. For instance,
is something called a normal subgroup of and hence from our library of group facts we know that is also a group, where that funny slash is called the “set quotient”, which your discrete mathematics teacher may have forced you to learn about. This fact turns out to be infinitely useful. For example, it helps us prove many things in number theory.
If you didn’t know what a group was, you would have produce obscure domain specific language to communicate these same ideas (as, for example, the founding fathers of number theory did, since they worked before groups). This would put a great mental burden on you, especially if you are interested in several subjects. Instead, we compress all this domain specific language into a single abstraction called a group and instead of saying “I employ this fact about object A that I spent 5 years proving” one instead says “This is a group and by standard facts about groups, such and such follows”. I think this is the practical sense in which abstraction is compression in mathematics, although I’m sure if you knew more information theory than I do you could make a nicer, more theoretical, argument.
Yes, I certainly agree that our compressor is currently more universally applicable, which is exactly why I am excited about seeing this sort of auto-compressing behavior in models like Mythos (and [now Sonnet 5 also](https://thezvi.substack.com/i/204364347/illegible-thinking-645)).
The basic idea, as I think of it, is this: in mathematics we have objects and we have facts about those objects. For example, if you consider the integers, , we have a pretty good sense of this object. There is an addition and a multiplication on , objects are invertible under addition (meaning for all there exists such that ), but not under multiplication. One can continue to list out facts. Now, is also a group under addition. Because is a group, it inherits all the facts that we know about groups. For instance,
is something called a normal subgroup of and hence from our library of group facts we know that is also a group, where that funny slash is called the “set quotient”, which your discrete mathematics teacher may have forced you to learn about. This fact turns out to be infinitely useful. For example, it helps us prove many things in number theory.
If you didn’t know what a group was, you would have produce obscure domain specific language to communicate these same ideas (as, for example, the founding fathers of number theory did, since they worked before groups). This would put a great mental burden on you, especially if you are interested in several subjects. Instead, we compress all this domain specific language into a single abstraction called a group and instead of saying “I employ this fact about object A that I spent 5 years proving” one instead says “This is a group and by standard facts about groups, such and such follows”. I think this is the practical sense in which abstraction is compression in mathematics, although I’m sure if you knew more information theory than I do you could make a nicer, more theoretical, argument.