The Descenders and the Absorbers: Two Perspectives on Deep Learning

Once upon a time there were two dissenting factions: the descenders and the absorbers. Both held incomplete but useful worldviews.

The descenders lived in the mountains. They had been approached by a prophet, who informed them that if they could reach the lowest valley in the land, they would achieve unbounded wisdom and knowledge. However, somewhat ironically, they had no map of the land, were completely blind, and had no internal sense of elevation.

To aid in this task, they were given a machine that could scan the landscape within a 5-foot radius around their current coordinate. However, each scan was partially corrupted and wrong. The descenders realized that if they changed the random seed of the scanner, it would corrupt in a different manner. The corruptions tended to cancel out, such that with enough seeds they could get a reasonably accurate picture of which way was downhill.

So they hatched a plan: run the scanner 1000 times, walk to the lowest elevation spot within the 5-foot radius, then repeat the process. Progress was steady but unbearably slow. 1000 scans took hours to complete, and 5 feet was almost no progress.

Over time they developed substantial optimizations. The first idea was to run only 10 scans, after which there were diminishing returns. The next finding was that even though the scan only covered 5 feet, they could typically step out at least 20.

However, sometimes the area of the scan had a misleading crevice and the extrapolation failed. So it was useful to scan several points in the neighborhood to cancel out the effect of misleading crevices. They stored the results of each scan in an archive, and used sophisticated algorithms to re-weight their step distance, scan count, and nearby sampling.

To the descenders, the ultimate challenge was to understand the topology of the land. Each scan was viewed as a corrupted and biased artifact, but necessary to power their compass. The goal was to reach a destination, and progress was viewed as moving downward.

The absorbers had a different view on the path to enlightenment. They were stewards of a library, and were on a mission to collect 100 books across every domain. Data was their prized possession. Initially they would go out and accept every random book they could find.

Some domains had many more books than others, and duplicates were popping up at increasing rates as their library filled. They found that the collection of new books in incomplete domains followed a power law, and progress was greatly slowing down. They theorized that with an optimal book curation process, progress would remain linear.

So they built a curriculum to fill the physics domain first, then biology, then fiction, then comics. However, the plan didn’t exactly work. First, it was quite challenging to find good sourcing from a single domain, and inefficient to pass up on books popping up in other domains. Books would wear out and go missing over time, necessitating a constant stream of new books in every domain. They settled on a middle ground, where they would still accept random books as they were found, but also made intentional efforts to target low-frequency domains.

Another challenge arose regarding the size of the library. They had not correctly anticipated the number of domains in the universe. Initially this resulted in turning away books while they constructed a larger library. Eventually they realized that the most efficient approach was to massively oversize their library up front, such that they rarely had to turn away a book. Even so, they instituted a duplicate policy to turn away books that were near-copies of books they already held.

To the absorbers, the ultimate challenge was the steady collection of knowledge. They paid little attention to the degrading books and books gone missing, and didn’t understand that dynamic fully, or the nature of compression. They focused heavily on having a large amount of space in their library, and searching out rare domains.

Both factions held useful but incomplete worldviews.

The main flaw in the descenders’ view was the number of degrees of freedom in their topology. They measured the location via latitude and longitude. But reality had countless dimensions, and could have further dimensions added beyond that. They also missed how each scan had useful knowledge beyond just the slope of the land. Yet they deeply understood the mechanics of descent.

The main flaw in the absorbers’ view was the rigidity of the book: in reality, books would go missing or degrade, and could be combined and compressed. Their model shifted too far towards rigid facts, and away from universal reasoning machinery. They also didn’t grasp that some books had a sort of non-linear topology at play, and couldn’t be accepted into the library in full until they were sampled and compressed many times. Yet they deeply understood the inefficiencies of procuring duplicate books and the value of library sizing.

Zooming out: if one of these two views sounds much closer to your view of deep learning than the other, then it is worth considering the alternative perspective more closely. The future of deep learning and understanding models may involve a combination of ideas, each of which is obvious to one perspective, but foreign to others.

No comments.