The big swing needs to engage directly with “what seems incentive compatible with ‘most people don’t care about truth, really’ and “pundits and leaders actively resist submitting to processes that could rule them out” and “any process that has teeth will be a target of politicization”.
I guess in that case. I consider myself engaged in that ‘big swing’. I am trying to create a norm of community notes and providence tracking on the web, both through my organisation (Goodheart Labs)′ work and through a fund I hope to launch soon focused on this work.
Personally I think community notes is a better big swing than forecasting.
But to engage with your theory.
you make it happen frictionless/automatically on a big platform people already use.
you make that happen by selling Musk on it (which seems hard but doable? modulo prototyping it in lower stakes places first and, like, you know, make a world-class UI and sort out novel algorithmic problems or whatever, but, “that’s the easy part”)
I think this is probably the hardest part
you use LLMs to automate it as much as you can (in particularly aggregating and summarizing the object level evidence).
you use community notes algorithm to handle “resolve complex predictions quickly without having to deal with resolution criteria, in a way that will hold up.”
I think this needs to happen at the start, not the end. There should be agreement while the forecast is in flight.
If you think this, do you think this would work on LessWrong, why why not? (I don’t think it would)
use various UI tricks to make it feel like a fun game instead of a punishment.
Like my sense is that people don’t like or aren’t capable of back rationalisation on controversial topics to come to conclusions they don’t like. But they are if they agreed beforehand. So personally I expect a process (sure, I agree, a community notes-like process) to happen during the forecast.
Personally I think community notes is a better big swing than forecasting.
Well insofar as community notes is the bedrock that enables this thing, yeah it seems basically “strictly better as a big swing.” You’d have to invent it first to get people warmed up to using it in other contexts.
Providence tracking is about knowing where images and videos were first seen. It is very useful as a way of trivially knowing that some claims backed up by images are inaccurate (because the image existed before the claimed time). Pangram for images is part of this, ie knowing which ones are AI generated.
I think this needs to happen at the start, not the end. There should be agreement while the forecast is in flight.
...
Like my sense is that people don’t like or aren’t capable of back rationalisation on controversial topics to come to conclusions they don’t like. But they are if they agreed beforehand. So personally I expect a process (sure, I agree, a community notes-like process) to happen during the forecast.
Mm, yeah sounds right. I assume here you mean agreement on the parameters/ironing-out-operationalization as much as is practical?
If you think this, do you think this would work on LessWrong, why why not? (I don’t think it would)
I feel something like “this is overkill for LessWrong.” (We’ve particularly tried out community-notes algorithms… and, well, for good or for ill, there just doesn’t actually seem to be major competing clusters in LW that need to be community-notes-algorithm’d – the results aren’t noticeably different from the normal karma system)
I think it’d be good to have various automated nudges on LW towards better epistemics. I think this isn’t “the piece that LW is missing” in a way that it feels more like “a piece that twitter is missing.” The reason to focus on predictions is because it’s a way to inject epistemics into the broader populace that it’s easier to justify.
Insofar as I believe this is the right thing for twitter, I do think I should prioritize making some version of it happen on LW and if it’s not working here that is evidence it wouldn’t work there.
Mm, yeah sounds right. I assume here you mean agreement on the parameters/ironing-out-operationalization as much as is practical?
And just what the rough thing actually means. Like sometimes a question turns out to mean quite different things. We can drill down on the distiction, but I guess there are examples where it isn’t res criteria, broadly construed.
I feel something like “this is overkill for LessWrong.” (We’ve particularly tried out community-notes algorithms… and, well, for good or for ill, there just doesn’t actually seem to be major competing clusters in LW that need to be community-notes-algorithm’d – the results aren’t noticeably different from the normal karma system)
I don’t think this is why you don’t have forecasts more widely locked/tracked on lesswrong.
Insofar as I believe this is the right thing for twitter, I do think I should prioritize making some version of it happen on LW and if it’s not working here that is evidence it wouldn’t work there.
I think a takeaway from this convo is actually “I want to build a singleplayer writing tool that just naturally prompts you to notice prediction-implications of your writing and help operationalize them”, and if that seems to be going well is more natural to try out in multiplayer contexts.
I don’t think this is why you don’t have forecasts more widely locked/tracked on lesswrong.
We’ve specifically thought about forecasts several times, and each time we end up feeling “idk, the forecasts just don’t quite seem to be doing the same type of work that LW posts are doing, like, the fact that you need to do all this annoying operationalization to make them trackable is pretty annoying and doesn’t really feel like it’s Doing The Thing.”
This thread has made me feel more optimistic about routing around that (less with community notes probably, more with AI assistance)
I guess in that case. I consider myself engaged in that ‘big swing’. I am trying to create a norm of community notes and providence tracking on the web, both through my organisation (Goodheart Labs)′ work and through a fund I hope to launch soon focused on this work.
Personally I think community notes is a better big swing than forecasting.
But to engage with your theory.
you make it happen frictionless/automatically on a big platform people already use.
you make that happen by selling Musk on it (which seems hard but doable? modulo prototyping it in lower stakes places first and, like, you know, make a world-class UI and sort out novel algorithmic problems or whatever, but, “that’s the easy part”)
I think this is probably the hardest part
you use LLMs to automate it as much as you can (in particularly aggregating and summarizing the object level evidence).
you use community notes algorithm to handle “resolve complex predictions quickly without having to deal with resolution criteria, in a way that will hold up.”
I think this needs to happen at the start, not the end. There should be agreement while the forecast is in flight.
If you think this, do you think this would work on LessWrong, why why not? (I don’t think it would)
use various UI tricks to make it feel like a fun game instead of a punishment.
Like my sense is that people don’t like or aren’t capable of back rationalisation on controversial topics to come to conclusions they don’t like. But they are if they agreed beforehand. So personally I expect a process (sure, I agree, a community notes-like process) to happen during the forecast.
Is that better?
Well insofar as community notes is the bedrock that enables this thing, yeah it seems basically “strictly better as a big swing.” You’d have to invent it first to get people warmed up to using it in other contexts.
Could you explain more about what that means?
Providence tracking is about knowing where images and videos were first seen. It is very useful as a way of trivially knowing that some claims backed up by images are inaccurate (because the image existed before the claimed time). Pangram for images is part of this, ie knowing which ones are AI generated.
Cool. I am glad you’re working on that!
I am not, but I hope to cause others to be. There is also a new standard.
Mm, yeah sounds right. I assume here you mean agreement on the parameters/ironing-out-operationalization as much as is practical?
I feel something like “this is overkill for LessWrong.” (We’ve particularly tried out community-notes algorithms… and, well, for good or for ill, there just doesn’t actually seem to be major competing clusters in LW that need to be community-notes-algorithm’d – the results aren’t noticeably different from the normal karma system)
I think it’d be good to have various automated nudges on LW towards better epistemics. I think this isn’t “the piece that LW is missing” in a way that it feels more like “a piece that twitter is missing.” The reason to focus on predictions is because it’s a way to inject epistemics into the broader populace that it’s easier to justify.
Insofar as I believe this is the right thing for twitter, I do think I should prioritize making some version of it happen on LW and if it’s not working here that is evidence it wouldn’t work there.
And just what the rough thing actually means. Like sometimes a question turns out to mean quite different things. We can drill down on the distiction, but I guess there are examples where it isn’t res criteria, broadly construed.
I don’t think this is why you don’t have forecasts more widely locked/tracked on lesswrong.
I would watch with interest.
I think a takeaway from this convo is actually “I want to build a singleplayer writing tool that just naturally prompts you to notice prediction-implications of your writing and help operationalize them”, and if that seems to be going well is more natural to try out in multiplayer contexts.
Interesting, I am interested.
We’ve specifically thought about forecasts several times, and each time we end up feeling “idk, the forecasts just don’t quite seem to be doing the same type of work that LW posts are doing, like, the fact that you need to do all this annoying operationalization to make them trackable is pretty annoying and doesn’t really feel like it’s Doing The Thing.”
This thread has made me feel more optimistic about routing around that (less with community notes probably, more with AI assistance)
Why do you think that’s different than twitter? Like of LW makes claims about the future, right?