I want the world to be more robust against small academic labs or startups building hyper efficient brain-like AGI / much better paradigm AI. Both during a pause or in LLMs-don’t-scale worlds.
Some thoughts:
Ideally there’s a widespread toolkit of evaluation, control techniques and model organism / elicitation approaches that are easy for the small lab to use. Potentially including being offered by external orgs.
Ideally it’s illegal to run experiments that could result in hyper efficient AIs with dangerous capabilities. Or at the least illegal to do so without alerting authorities and using safety toolkits. A lot of academic labs probably pay attention to what’s legal.
Maybe there should be easy to use channels of transferring algorithmic breakthroughs or potentially just a model API to frontier labs who have better control, automated alignment research scaffolds and d/acc programs.
The whole thing plausibly looks like having a bunch of open research on brain algorithms and AI already published with a few tweaks and optimizations left before it’s clear the new paradigm gets to dangerous capabilities. Ideally that research would have already been illegal before triggering noticeable capabilities gains, that seems hard to enforce though, but if ppl take the problem seriously after some future wake up then maybe it’d be possible.
There should probably be some tracking of groups doing research relevant to hyper efficient / brain-like AGI and social pressure on them to be careful or stop and ideally monitoring by possible future sane-after-wake-up intelligence agencies so long as the tracking doesn’t make it more salient what the algorithmic insights are and they are fine tuned and developed much earlier and less safely because of it.
I want the world to be more robust against small academic labs or startups building hyper efficient brain-like AGI / much better paradigm AI. Both during a pause or in LLMs-don’t-scale worlds.
Some thoughts:
Ideally there’s a widespread toolkit of evaluation, control techniques and model organism / elicitation approaches that are easy for the small lab to use. Potentially including being offered by external orgs.
Ideally it’s illegal to run experiments that could result in hyper efficient AIs with dangerous capabilities. Or at the least illegal to do so without alerting authorities and using safety toolkits. A lot of academic labs probably pay attention to what’s legal.
Maybe there should be easy to use channels of transferring algorithmic breakthroughs or potentially just a model API to frontier labs who have better control, automated alignment research scaffolds and d/acc programs.
The whole thing plausibly looks like having a bunch of open research on brain algorithms and AI already published with a few tweaks and optimizations left before it’s clear the new paradigm gets to dangerous capabilities. Ideally that research would have already been illegal before triggering noticeable capabilities gains, that seems hard to enforce though, but if ppl take the problem seriously after some future wake up then maybe it’d be possible.
There should probably be some tracking of groups doing research relevant to hyper efficient / brain-like AGI and social pressure on them to be careful or stop and ideally monitoring by possible future sane-after-wake-up intelligence agencies so long as the tracking doesn’t make it more salient what the algorithmic insights are and they are fine tuned and developed much earlier and less safely because of it.