Part of my answer to this is “we have explained our models in detail several times, and I haven’t seen you reference any of those details in your suggestions. You can’t propose a workable idea until you have actually integrated everything we’ve said.”
Our most valuable resource is habryka’s time. We’re happy to talk things through, but by now I think habryka has explained all the pieces of the problem at least twice.
I think things will go better here if you start by rereading everything (at least from this thread, and maybe any specific things he linked to from the past couple rounds of discussion), starting with the assumption “we are not doing this because Eliezer made us, we are doing this because we think it is a good idea”, and try to summarize everything as you understand it so far.
In this case I actually did feed the thread to Gemini 3.1 Pro and asked it how habryka would answer my question, and it said 1 or 4, and then habryka’s actual answer turned out to be closer to 2. Its reasoning, which I thought was plausible, was that habryka has a vision of the archipelago that he wants to achieve, that depends on authors being mods, so the cost of site mods wasn’t really a crux for him. So I was kind of surprised when I saw habryka’s actual answer.
I suspect you may be overestimating how clear your writings/answers are, or how much you explained things, or how relevant previous explanations are given new assumptions, or how much I can/should trust you/others to represent habrya. For example, I asked earlier “Why can’t one of the other experienced mods train or manage the new mod?” and got no answer to this (confirmed via LLM that it wasn’t answered or explained anywhere in the entire comment tree rooted at the shortform OP), so there’s a big blank spot in my model of what’s going on at Lightcone.
Ah, well cool that you tried that and sorry it didn’t work.
But to clarify, I would be using Fable for this, I don’t have a belief that older models are able to track the arguments here.
I don’t have that strong a belief that this “ask the models” thing will turn out to work, maybe that’s a dead-end for now. But, insofar as the idea has promise I’d be doing this with Fable, making sure it’s read Meta-tations on Moderation: Towards Public Archipelago , Banning Said Achmiz (and broader thoughts on moderation) and all comments on those, and reading your shortform page for all comments between you and habryka on moderation.
One reason I think this problem is on your side is that Oliver has said multiple times “we are not doing this because Eliezer said so”, and you repeated the belief that we were in your most recent interaction.
There are less straightforward reasons for that but if you’re still ignoring basic statements like that the conversation feels pretty hopeless to me.
For example, I asked earlier “Why can’t one of the other experienced mods train or manage the new mod?” and got no answer to this (confirmed via LLM that it wasn’t answered or explained anywhere in the entire comment tree rooted at the shortform OP)
Oh I see. To explain why I “ignored” it, I thought what you wrote wasn’t directly relevant to what I was trying to do at the time, which was to test whether Lightcone might be amenable to my offer. I had my reasons for thinking that the offer would be a win for me despite what you wrote, basically because I think you didn’t quite understand how I thought it would solve (or at least possibly solve) the problem, but I didn’t feel like taking the time to explain. Because if Lightcone wouldn’t be amenable to the offer, then explaining my model of how it would help me would be kind of pointless.
(i.e. basically, @Wei Dai you can have Claude Code Fable scrape for all the past moderation discussions, and then write up a message and ask “are there replies that habryka pretty obviously say to this?” and then iterate a bit until you get to something that feels like it’s at least moving the conversation forward)
(It’s not totally obvious how to do the scraping part; maybe you ask CC to scrape all content from a list of users you give it, and then ask it to search / read for everything about moderation and put together a timeline of all that content, and then interview it about that blob?)
I’ve done this actually for twitter conversations with people I don’t know and it was just pretty straightforward. “Hey I want to talk to so-and-so about X, please read everything by so-and-so that bears on X”.
One important thing was to fork the conversation after the initial “read up everything relevant”, so that when I asked “okay, what if I wrote this to them?”, it doesn’t anchor on whatever random stuff it guessed the first couple times.
The failure mode I’m concerned about is missing big chunks of content, so you only get 1⁄2 of the relevant stuff or something. It may not be that bad a failure mode, but like, for example, just asking a Fable chat probably would turn up far under 90% of the past convos between Wei / Zack / LW mods? CC would probably do better? (I have done a significant amount of using CC to “read” many papers, and I have to do nonzero poking CC to improve fetching methods; by default it gives up for various reasons even when there is a way.)
I agree it might not be totally comprehensive, but, that seems kinda fine? A lot of the concepts have gotten repeated. (Like, seems fine to try a lower effort version and see how it goes and try a higher effort version if it feels on-track-to-be-helpful but insufficient)
There are some difficulties with that because of system prompts and default tools.
Claude.ai/ClaudeCode by default read websites using the built-in WebFetch tool. The WebFetch tool loads the HTML, turns it into Markdown and feeds it into Haiku(! Edit: Claude.ai’s tool does not feed it into Haiku), which summarizes (with strong instructions for limiting verbatim quotes!) and only that output goes to Fable. (By contrast ChatGPT/Codex’s default tool gives the full text, but also has strong instructions for limiting verbatim quotes. For Codex this is only in the tool instruction.)
ClaudeCode/Codex can easily work around this by using e.g. cURL (especially if you tell them about lesswrong.com/api). Edit: Claude.ai can too, if you enable Network egress in the settings. Both can of course create you a re-usable skill. And with some prompting you can get around the quoting limitations.
Right, I’ve had cc download my whole LW history, and I had to make a system to actually download the actual entire whole paper. I think one can easily do it by asking cc, but one has to specifically ask.
Part of my answer to this is “we have explained our models in detail several times, and I haven’t seen you reference any of those details in your suggestions. You can’t propose a workable idea until you have actually integrated everything we’ve said.”
Our most valuable resource is habryka’s time. We’re happy to talk things through, but by now I think habryka has explained all the pieces of the problem at least twice.
I think things will go better here if you start by rereading everything (at least from this thread, and maybe any specific things he linked to from the past couple rounds of discussion), starting with the assumption “we are not doing this because Eliezer made us, we are doing this because we think it is a good idea”, and try to summarize everything as you understand it so far.
In this case I actually did feed the thread to Gemini 3.1 Pro and asked it how habryka would answer my question, and it said 1 or 4, and then habryka’s actual answer turned out to be closer to 2. Its reasoning, which I thought was plausible, was that habryka has a vision of the archipelago that he wants to achieve, that depends on authors being mods, so the cost of site mods wasn’t really a crux for him. So I was kind of surprised when I saw habryka’s actual answer.
I suspect you may be overestimating how clear your writings/answers are, or how much you explained things, or how relevant previous explanations are given new assumptions, or how much I can/should trust you/others to represent habrya. For example, I asked earlier “Why can’t one of the other experienced mods train or manage the new mod?” and got no answer to this (confirmed via LLM that it wasn’t answered or explained anywhere in the entire comment tree rooted at the shortform OP), so there’s a big blank spot in my model of what’s going on at Lightcone.
Ah, well cool that you tried that and sorry it didn’t work.
But to clarify, I would be using Fable for this, I don’t have a belief that older models are able to track the arguments here.
I don’t have that strong a belief that this “ask the models” thing will turn out to work, maybe that’s a dead-end for now. But, insofar as the idea has promise I’d be doing this with Fable, making sure it’s read Meta-tations on Moderation: Towards Public Archipelago , Banning Said Achmiz (and broader thoughts on moderation) and all comments on those, and reading your shortform page for all comments between you and habryka on moderation.
One reason I think this problem is on your side is that Oliver has said multiple times “we are not doing this because Eliezer said so”, and you repeated the belief that we were in your most recent interaction.
There are less straightforward reasons for that but if you’re still ignoring basic statements like that the conversation feels pretty hopeless to me.
fyi I have stopped answering questions like this after you ignored most of the content of this comment, which stated pretty clearly what sorts of discussion would be likely to move us, and why: https://www.lesswrong.com/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform?commentId=betXXw7gFskNfiHp3
Oh I see. To explain why I “ignored” it, I thought what you wrote wasn’t directly relevant to what I was trying to do at the time, which was to test whether Lightcone might be amenable to my offer. I had my reasons for thinking that the offer would be a win for me despite what you wrote, basically because I think you didn’t quite understand how I thought it would solve (or at least possibly solve) the problem, but I didn’t feel like taking the time to explain. Because if Lightcone wouldn’t be amenable to the offer, then explaining my model of how it would help me would be kind of pointless.
(Cf. https://www.lesswrong.com/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform?commentId=DuiewixJiufzTH84h )
Oh that is pretty good actually.
(i.e. basically, @Wei Dai you can have Claude Code Fable scrape for all the past moderation discussions, and then write up a message and ask “are there replies that habryka pretty obviously say to this?” and then iterate a bit until you get to something that feels like it’s at least moving the conversation forward)
(It’s not totally obvious how to do the scraping part; maybe you ask CC to scrape all content from a list of users you give it, and then ask it to search / read for everything about moderation and put together a timeline of all that content, and then interview it about that blob?)
I’ve done this actually for twitter conversations with people I don’t know and it was just pretty straightforward. “Hey I want to talk to so-and-so about X, please read everything by so-and-so that bears on X”.
One important thing was to fork the conversation after the initial “read up everything relevant”, so that when I asked “okay, what if I wrote this to them?”, it doesn’t anchor on whatever random stuff it guessed the first couple times.
The failure mode I’m concerned about is missing big chunks of content, so you only get 1⁄2 of the relevant stuff or something. It may not be that bad a failure mode, but like, for example, just asking a Fable chat probably would turn up far under 90% of the past convos between Wei / Zack / LW mods? CC would probably do better? (I have done a significant amount of using CC to “read” many papers, and I have to do nonzero poking CC to improve fetching methods; by default it gives up for various reasons even when there is a way.)
I agree it might not be totally comprehensive, but, that seems kinda fine? A lot of the concepts have gotten repeated. (Like, seems fine to try a lower effort version and see how it goes and try a higher effort version if it feels on-track-to-be-helpful but insufficient)
There are some difficulties with that because of system prompts and default tools.
Claude.ai/ClaudeCode by default read websites using the built-in WebFetch tool. The WebFetch tool loads the HTML, turns it into Markdown and feeds it into Haiku(! Edit: Claude.ai’s tool does not feed it into Haiku), which summarizes (with strong instructions for limiting verbatim quotes!) and only that output goes to Fable. (By contrast ChatGPT/Codex’s default tool gives the full text, but also has strong instructions for limiting verbatim quotes. For Codex this is only in the tool instruction.)
ClaudeCode/Codex can easily work around this by using e.g. cURL (especially if you tell them about lesswrong.com/api). Edit: Claude.ai can too, if you enable Network egress in the settings. Both can of course create you a re-usable skill. And with some prompting you can get around the quoting limitations.
Right, I’ve had cc download my whole LW history, and I had to make a system to actually download the actual entire whole paper. I think one can easily do it by asking cc, but one has to specifically ask.