dynomight
Not AI psychosis exactly, but I know a lot of people who get access to tools like Claude Code and start to feel the potential, the amazing potential and can’t seem to think about anything other than building their amazing new strongly-typed functional meta-agent Zettelkasten system, and how they’ve finally figured it out, the ultimate organization of all knowledge, and how it will change their life and the world, but there is no externally visible output other than an increase in them talking about said strongly-typed functional meta-agent Zettelkasten system.
I’ve thought about this a lot and I eventually decided the best solution was to “react more”. For whatever reason, I’ve always thought of reactions as kind of cringe. But I think they’re the best way we have to fight this bias. I’ve come around to Less Wrong’s extremely extensive list of reactions and I think we should all consider intentionally trying to increase our use of them.
BTW, I wish I had clarified: I’m not talking about any of the detailed critiques that people have made, e.g. any of the critiques mentioned here: https://www.lesswrong.com/posts/RPgHythvMKh6eG9pS/our-response-to-seb-krier-on-plan-a This was more inspired by a lot of quick comments I’ve seen that seem to dismiss Plan A at a high level without really engaging with the details much at all. I’ve seen these in lots of places but I don’t feel comfortable linking to them because they’re in a sort of conversational context and it seems norm-breaking to single out individual people for just writing a few sentences.
Reading some of the critiques of Plan A online, I’m increasingly convinced that many people start with a premise that we live in a “normal” world, where only “normal” things happen. With that premise, anything that predicts that wild and crazy things could happen must be wrong. I think that premise is clearly empirically false, but it’s hard to see because all the previous wild and crazy reality-shifting things that have already happened are now accepted as “normal”.
Two brief random thoughts:
-
I think it’s great how large a role robotics plays in this analysis. In general, I feel like robotics isn’t given nearly the amount of attention it deserves. Without robotics, AI’s influence on the world is bottlenecked by human hands. So the speed of the feedback loop between AI building better robots to build better robot factories to build better robots seems like a very important crux. So bravo for taking that issue seriously.
-
Regarding the proposal to make AI research public but not model weights, I think it might be worth giving a bit more thought to the idea the whatever is made public essentially becomes a public good, and public goods tend to be under-supplied—or in this case perhaps “under-supplied”, as reducing the amount of AI research supplied might be considered a good thing or perhaps even the primary benefit of such a policy. Theoretically, if people really were forced to make all AI research public in a way that was instantly useful to competitors, it seems like that would either drastically reduce investment in research or force people to invest in research that other people couldn’t take advantage of, e.g. research that is only useful given certain propriety data or hardware or something. I understand the logic for not making weights public, namely it would be hard to get people to agree, plus you can easily tune out safeguards. But on the flip side, it seems very hard to formalize what exactly counts as “research” (particularly if everyone is trying to skirt that line). Radical total transparency is definitely a big ask, but it seems easier to enforce if agreed. And if weights were mandated to be public, that might also reduce investment, which would be good. Or, again, it might incentivize people to build models where other people couldn’t make use of the weights with out some kind of propriety hardware. (To be clear, I still lean against public weights for the reasons given in the proposal. But I’d be interested to hear a more discussion of this.)
-
I guess the general argument is that the optimization space is very high-dimensional, and there’s no pressure to reason in English, so it just seems incredibly unlikely that the optimum would look like thinking aloud in English. I’m not sure how much I buy the human analogy; we have different capabilities and IMO humans don’t think in language to the same degree that current LLMs do. But certainly humans in specialized areas find it useful to invent lots of specialized terminology and conventions. If you sent people to an island and had them live for 10000 years with their society entirely devoted to chess, I’d conjecture that their language would evolve quite a bit. If RL could “really do RL” it should be able to create such evolutions for many domains, or (I suspect) cross-domain.
DeepSeek (which GRPO was invented for) still uses a normal English CoT to this day. (V4 Pro) They had to alternate with supervised learning because RL is just very very hard and GRPO is incredibly weak compared to supervised learning. They aren’t deliberately sacrificing performance by forcing an English CoT instead of a more useful or efficient neuralese. They are forced to do it because RL just doesn’t work that well.
Well, if RL “really worked”, yes it seems like it should almost certainly invent and reason in an alien language. But with all the models where we can actually see the CoT, it’s more or less in English. My interpretation (and I think the standard interpretation?) for that is that RL is hard and can’t really optimize too far from the parameters found in pretraining.
Thanks, so is the interpretation something like this?
2a. “this is what Fable’s CoT looks like, but it doesn’t follow that we should accelerate our timelines, because we already knew RL could produce this kind of behavior, so it should already be baked into your timelines if your timelines are good.”
I don’t mind disagree votes, I’m just interested to see what people think. I was very surprised to see that CoT, but I’m only really basing that on reading CoT from open models (Kimi, DeepSeek, etc.) that do sometimes have random meltdowns, but AFAIK never do something like this productively.
Perfect, thanks, just what I was looking for.
An additional unclear thing about HIIT is that we don’t really know with confidence that VO2max causally improves the things we “really” care about. I mean… as you say, any exercise that improves VO2max almost certainly improves health (a lot), but it’s entirely possible that exercise program A might improve VO2max slightly less than exercise program B and yet exercise program A might be better for health.
(Totally endorse your logic in the last paragraph BTW. I think, if anything, this strengthens that argument?)
Not sure how to interpret the accuracy downvotes. :) Do they mean:
The explanation is something other than “random meltdown” or “this is what Fable’s CoT looks like”?
The explanation is “this is what Fable’s CoT looks like, but it doesn’t follow that we should accelerate our timelines”?
Something else?
Has there been any discussion here of the “leaked Fable chain of thought”?
In particular, does anyone have any view on if this is just a random meltdown or if this is actually what Fable’s CoT looks like? If the latter, perhaps that should substantially accelerate our timelines? (On the logic that reinforcement learning can accomplish more than was understood.)
This is great, but I’d caution against saying it’s definitive. There’s a risk of multiple hypothesis testing, as well as the usual risks of publication bias, possible errors in the way individual studies were conducted, etc.
Yeah, I totally agree. It’s entirely possible that the vitamin-D / folate hypothesis is true, but the benefits of vitamin D come in ways that wouldn’t impact all-cause-mortality. I sort of hinted at things along this line with “even if true, might be driven by severe deficiency and rickets” but maybe I could do a better job of reflecting that there could be other things on that list.
With the result of the Bores campaign out, here’s something I’m wondering: Suppose someone has good AI positions and is thinking about running for office. How easy would it be for them to tap into the same pool of support that was behind Bores? Would it be easy? And, if it is easy, how legible is that easiness? Could some structures in place that would make it easier / legible-er? If that were done, would that significantly move the needle on how likely such people are to run for office?
I suppose the pessimistic take is that you don’t want this to be too easy or legible, lest you encourage opportunists. But it seems like some existing organizations have a process that is fairly legible (e.g. the NRA)?
I’d hedge a little. I think it’s fair to say that doctors and scientists have proven that severe deficiency (below ~25 nmol/L) is definitely bad. The question is if going from 25 to 40 or 40 to 60 nmol/L is helpful. Personally, I think it’s genuinely possible that the answer is no, but in the absence of certainty, it seems pretty clear that the safest thing is to avoid those low levels.
Statistically speaking, I 100% agree with you. The primary reason these trials didn’t give us a clear answer is that they weren’t testing “the benefits of supplementing vitamin D if you’re low” but more something like “the benefits of giving extra vitamin D to a population where most people are already supplementing”.
OK, more-boring mistakes form installed.
Thank you! I didn’t look much at studies looking at the impact of vitamin D on mood, but at a glance, the RCTs seem incredibly inconclusive. This meta analysis found a good looking effect-size, but the funnel plot suggests a horrific amount of publication bias. There’s an ongoing study that looks well-designed but it’s happening in Finland (which already fortifies their food to the sky) and they don’t screen for low baseline levels and the sample is small, so it seems unlikely to be definitive...
I’m a little bit out of the loop… How and why would you serve markdown content for LLMs?