it doesn’t even work that well for capabilities anymore.
Say more about this? I thought it still worked well, modulo the caveat that you now need reasonably good data scaling (and good RL environments, which is why data spending is rising fast), or is that the caveat that makes assuming magic ood generalization not work nearly as well anymore than previously?
Say more about this? I thought it still worked well, modulo the caveat that you now need reasonably good data scaling (and good RL environments, which is why data spending is rising fast), or is that the caveat that makes assuming magic ood generalization not work nearly as well anymore than previously?
modern capabilities stacks are very complex
Isn’t that mostly for “specific capabilities benefit from specific RL, and RL is as treacherous as ever” reasons? Very dependent on how you train?