https://verifierchallenge.org/
Nice, I’m imagining something a lot larger and higher profile ideally, but I like this and will reach out to them.
https://verifierchallenge.org/
Nice, I’m imagining something a lot larger and higher profile ideally, but I like this and will reach out to them.
Interesting! I would say I get moderate uplift from Claude (Code, with a custom skill providing lots of context about my work, and style guidance) but that it is still notably worse than getting a review by a smart high-context human.
Maybe try sending me various docs that you are comfortable sharing and I will run them through my claude code skill and share the outputs with you?
Why does this post say ’29 min read, 7,100 words’ - could it be something to do with the embedded interactive elements, maybe the code for those is automatically counted by the reading time estimator? Minor, but seems worth fixing if the bug occurs elsewhere too.
Only if we are better at moral reflection than Claude! It seems quite possible to me that Claude 7 can better CEV my values than I can, and that its ethical maturity is ‘better’ than mine in some sense.
But I agree that there is an important cost to making AIs less corrigible.
I liked https://firstscattering.com/p/red-lines-for-recursive-self-improvement as a quick initial discussion of possible places to draw a line.
I only read the LW version not the paper, but this seems like important work to me and I’m glad you’re doing it! What did you make of these two recent papers?
I have done some work on the policy side of this (whether we should/how we could enforce CoT monitorability on AI developers, or at least gain transparency into how monitorable SOTA models are). Lmk if ever it would be useful to talk about that, otherwise I will be keen to see where this line of work ends up!
I’d be interested in anyone’s thoughts on when to use this vs e.g., METR’s time horizon. The latter is of course more coding-focused than this general-purpose compilation, but that might be a feature not a bug for our purposes (predicting takeoff).
AI direction could make most workers much closer in productivity to the best workers. The difference between the productivity of the average and the best manual workers is perhaps around 2-6X
Based on the derivation, it seems you mean the difference in productivity of workers doing similar tasks in the same industry, which seems important to specify. Otherwise as written, I would say the “difference between the productivity of the average and the best manual workers” is >1000x between e.g. surgeons in rich countries and e.g. farm hands/construction workers/salespeople, etc in poor countries.
But it’s not clear to me the relevant multiplier is the one you pick within one country and industry. E.g. if we have abundant cheap AI cognitive labour, couldn’t I set up a company producing widgets in e.g. India, employ heaps of low-skill workers for cheap but make them very productive with AI training and direction, and make a killing?
Maybe the bottleneck here is more on political economy and insitution quality, such that even with AGI not all poor countries suddenly become rich because they have productive AI-led firms.
Overall I feel a bit confused how big I think the one-time boost would be, but if we are counting across countries I would suspect >10x. Perhaps in practice the US (or whoever has the intelligence explosion) would limit access to cognitive abundance to itself and maybe a few allies.
Great question, I don’t have deep technical knowledge here, but would also be very curious about this. Intuitively, that seems right that CoT monitoring doesn’t transfer over very well to this case.
Nice!
For the 2024 prediction “So, the most compute spent on a single training run is something like 5x10^25 FLOPs.” you cite v3 as having been trained on 3.5e24 FLOP, but that is outside an OOM. Whereas Grok-2 was trained in 2024 with 3e25, so seems to be a better model to cite?
I will note the rationalist and EA communities ahve committed multiple ideological murders
Substantiate? I down- and disagree-voted because of this un-evidenced very grave accusation.
I think I agree with your original statement now. It still feels slightly misleading though, as while ‘keeping up with the competition’ won’t provide the motivation (as there putatively is no competition), there will still be strong incentives to sell at any capability level. (And as you say this may be overcome by an even stronger incentive to hoard frontier intelligence for their own R&D and strategising use. But this outweighs rather than annuls the direct economic incentive to make a packet of money by selling access to your latest system.)
I think some analogies to evolutionary biology and ecology could be useful here. This result is about within one population the CDT allele being driven (eventually, not observed in these experiments) to fixation. But suppose we start with many separate populations, and then the populations with highest average fitness over time colonise more of the ecosystem. Or in this case you start with Kimi, Qwen, Gemma, etc, and then the ones that are most likely to cooperate lead to higher scores for their group, and that group becomes more common in the overall meta-population. And that effect should favour EDT.
Are there theoretical reasons to think the within group or between group selection effects will be stronger?