Seeing the kind of utter bullshit that ANNs tolerate routinely makes me kind of bullish on “imperfect uploads”.
If ANNs can degrade gracefully on perturbations like pruning, quantization, noise injection, model surgeries and more, then, what does that tell us of the robustness of biological NNs—networks that are, by their very nature, optimized to run in noisier, less perfect environments than that of deterministic silicon?
From what I’ve seen on the biological end, there are also hints that brains have their own scaling patterns—and the more “complex” you go, the more of the overall behavior is driven by topology—local and global connectivity—rather than hardwired specialized behaviors of singular neurons.
When you run 100 neurons total, each neuron is a specialized unit doing a specific thing, and it’s absolutely vital to get the specific neurons right to recreate behavior. When you run 100 000 neurons, neurons themselves become far more generic and interchangeable, and the behavior becomes far more connectome-driven. “Identify every single neuron type and characterize the behavior of each type extensively” is vital on one end of the spectrum, but may be “extra credit” on the other. Unprincipled “take 120 pre-made neuron models and brute force through them to find the combinations that seem to fit a few recorded patterns best” might get most of the way there, and much faster.
It seems likely that this trend would continue onwards, into millions and billions. Which bodes very well for those more connectome-centric “assume simplified neurons” approaches. More so when paired with the likely perturbation resistance.
I look favorably at the “don’t chase perfection, chase integration and scale” approach in this demo because of it. I get why it’s controversial—I just think the tradeoffs they made are quite sensible. Demoing obviously imperfect and incomplete but “good enough that it looks biologically plausible” behavior in a sim beats going for perfection a decade down the line, in my eyes. And the field does deserve more attention than it’s getting.
It does seem likely that bio brains are pretty robust to perturbation, but quantization produces mostly-independent noise. a structural difference across the entire model can produce potentially large systemic behavior differences. it only takes maybe 1ug lsd in the brain (out of a 100ug oral dose) to amplify into a huge difference. I asked claude to estimate and was told that’s an average of 10 molecules per synapse! a small out-of-distribution signaling difference and the entire thing is in a different kind of attractor. if you can be sure your lossiness is independent, you’re more likely to make use of the incredible redundancy. if you don’t know about a critical signaling pathway, then everything works within some regime and then breaks as soon as that signaling pathway is hit. and you have to be able to detect when that signaling pathway was important to know if you succeeded. Which means you need some sort of high bandwidth inspection to see if the dynamics are different under the conditions you care about.
also, there are known to be some systems in large brains that depend on relatively few neurons
This seems to me like a soft constraint problem at its core.
“How many little pins do you need to pin down the behavior of this system and force it into some subspace that’s reasonably close to the solution?”
I.e. for the LSD example: it’s not a “large scale” phenomenon. LSD’s effects are already prominent in tiny neuron populations “in vitro”. It’s not a “tiny almost imperceptible error that compounds to become catastrophic at scale”—it’s “already a big error that simply scales upwards well”.
Which means: you can have a “tiny neuron population in vitro” reference, and that alone will expose the “we made every single neuron in the brain sim act like it’s on LSD” failure mode. Many others like it too.
How many little pins? How many imposed constraints does it take to bleed off the bulk of major systematic failures? How many constraints does it take to walk away from “catastrophic instability” and enter the “mostly convergent” subspace, where innate perturbation resistance begins to work in our favor? How many of those constraints can be imposed with relatively straightforward means, long before you need to collect large bodies of new data and build specialized tooling like high bandwidth inspection?
The only real way to figure out is to go for it, and see what works and what fails. It’s an empirical problem, there’s no way around it. You can only test your assumptions if you make them first.
Seeing the kind of utter bullshit that ANNs tolerate routinely makes me kind of bullish on “imperfect uploads”.
If ANNs can degrade gracefully on perturbations like pruning, quantization, noise injection, model surgeries and more, then, what does that tell us of the robustness of biological NNs—networks that are, by their very nature, optimized to run in noisier, less perfect environments than that of deterministic silicon?
From what I’ve seen on the biological end, there are also hints that brains have their own scaling patterns—and the more “complex” you go, the more of the overall behavior is driven by topology—local and global connectivity—rather than hardwired specialized behaviors of singular neurons.
When you run 100 neurons total, each neuron is a specialized unit doing a specific thing, and it’s absolutely vital to get the specific neurons right to recreate behavior. When you run 100 000 neurons, neurons themselves become far more generic and interchangeable, and the behavior becomes far more connectome-driven. “Identify every single neuron type and characterize the behavior of each type extensively” is vital on one end of the spectrum, but may be “extra credit” on the other. Unprincipled “take 120 pre-made neuron models and brute force through them to find the combinations that seem to fit a few recorded patterns best” might get most of the way there, and much faster.
It seems likely that this trend would continue onwards, into millions and billions. Which bodes very well for those more connectome-centric “assume simplified neurons” approaches. More so when paired with the likely perturbation resistance.
I look favorably at the “don’t chase perfection, chase integration and scale” approach in this demo because of it. I get why it’s controversial—I just think the tradeoffs they made are quite sensible. Demoing obviously imperfect and incomplete but “good enough that it looks biologically plausible” behavior in a sim beats going for perfection a decade down the line, in my eyes. And the field does deserve more attention than it’s getting.
It does seem likely that bio brains are pretty robust to perturbation, but quantization produces mostly-independent noise. a structural difference across the entire model can produce potentially large systemic behavior differences. it only takes maybe 1ug lsd in the brain (out of a 100ug oral dose) to amplify into a huge difference. I asked claude to estimate and was told that’s an average of 10 molecules per synapse! a small out-of-distribution signaling difference and the entire thing is in a different kind of attractor. if you can be sure your lossiness is independent, you’re more likely to make use of the incredible redundancy. if you don’t know about a critical signaling pathway, then everything works within some regime and then breaks as soon as that signaling pathway is hit. and you have to be able to detect when that signaling pathway was important to know if you succeeded. Which means you need some sort of high bandwidth inspection to see if the dynamics are different under the conditions you care about.
also, there are known to be some systems in large brains that depend on relatively few neurons
This seems to me like a soft constraint problem at its core.
“How many little pins do you need to pin down the behavior of this system and force it into some subspace that’s reasonably close to the solution?”
I.e. for the LSD example: it’s not a “large scale” phenomenon. LSD’s effects are already prominent in tiny neuron populations “in vitro”. It’s not a “tiny almost imperceptible error that compounds to become catastrophic at scale”—it’s “already a big error that simply scales upwards well”.
Which means: you can have a “tiny neuron population in vitro” reference, and that alone will expose the “we made every single neuron in the brain sim act like it’s on LSD” failure mode. Many others like it too.
How many little pins? How many imposed constraints does it take to bleed off the bulk of major systematic failures? How many constraints does it take to walk away from “catastrophic instability” and enter the “mostly convergent” subspace, where innate perturbation resistance begins to work in our favor? How many of those constraints can be imposed with relatively straightforward means, long before you need to collect large bodies of new data and build specialized tooling like high bandwidth inspection?
The only real way to figure out is to go for it, and see what works and what fails. It’s an empirical problem, there’s no way around it. You can only test your assumptions if you make them first.