I very much agree but with even wider uncertainty. Average verbosity could probably be anywhere from 1.25x to 2.5x depending on methodology, setting, etc and it would really narrow down our uncertainty if we knew.
One simple experiment to estimate verbosity would be to get a large sample of AI-written and human-written code, and tell AIs to delete as many useless lines of code in each as possible while preserving functionality. This gives us the average verbosity factor of AI vs human written code. If it’s high one could maybe rule out >3x research uplift already. Then we’d want to split these datasets by production vs research code, and after that use LLM judges to estimate the amount of barely useful code. There’s a huge amount of low hanging fruit here.
I very much agree but with even wider uncertainty. Average verbosity could probably be anywhere from 1.25x to 2.5x depending on methodology, setting, etc and it would really narrow down our uncertainty if we knew.
One simple experiment to estimate verbosity would be to get a large sample of AI-written and human-written code, and tell AIs to delete as many useless lines of code in each as possible while preserving functionality. This gives us the average verbosity factor of AI vs human written code. If it’s high one could maybe rule out >3x research uplift already. Then we’d want to split these datasets by production vs research code, and after that use LLM judges to estimate the amount of barely useful code. There’s a huge amount of low hanging fruit here.