Deep learning understood as a process of up- and down-weighting circuits is incredibly similar conceptually to logical induction.
Pre- and post-training LLMs is like juicing the market so that all the wealthy traders are different human personas, then giving extra liquidity to the ones we want.
I expect that the process of an agent cohering from a set of drives into a single thing is similar to the process of a predictor inferring the (simplicity-weighted) goals of an agent by observing it. RLVR is like rewarding traders which successfully predict what an agent which gets high reward would do.
Logical Induction doesn’t get you all the way, since the circuits can influence other circuits, like traders that are allowed to bet on each other, or something.
(These analogies aren’t quite perfect, I swapped between trading day-as-training batch and trading day-as-token)
Somebody must have made these observations before, but I’ve never seen them.
Spitballing:
Deep learning understood as a process of up- and down-weighting circuits is incredibly similar conceptually to logical induction.
Pre- and post-training LLMs is like juicing the market so that all the wealthy traders are different human personas, then giving extra liquidity to the ones we want.
I expect that the process of an agent cohering from a set of drives into a single thing is similar to the process of a predictor inferring the (simplicity-weighted) goals of an agent by observing it. RLVR is like rewarding traders which successfully predict what an agent which gets high reward would do.
Logical Induction doesn’t get you all the way, since the circuits can influence other circuits, like traders that are allowed to bet on each other, or something.
(These analogies aren’t quite perfect, I swapped between trading day-as-training batch and trading day-as-token)
Somebody must have made these observations before, but I’ve never seen them.