I suspect that the OP meant something like @Richard_Ngo’s rants about how EAs’ philosophy contributed to the AI race without a profound understanding of the AIs. IIRC some others claimed that EAs’ views did stuff like diaincentivizing Kelly betting and failing to understand until the ChatGPT moment why the AGI was due by 2030 or earlier, not 2050.
StanislavKrym
I wonder if this is, to some extent, an artifact of the AA index. IIRC there was a study on LW which found that the benchmarks comprising the ECI have an extra direction of agentism vs math and Claudes’ capabilities were skewed in that very direction and GPTs’ models were skewed in the other direction.
EpochAI’s ECI measurements had GPTs overtake Claudes until GPT-5.5 Pro (ECI 162) was compared to Fable 5 (ECI 164). Then, in July, OAI responded with GPT-5.6 Sol, which has Opus’ prices and an ECI of 162, Anthropic released Opus 5 with the ECI of 163, and in early September Anthropic and OAI released Fable 5.1 and GPT-6 Astra which have the ECIs of 165 and 167. Alas, I don’t understand which capabilities are most important for a potential FOOM.
@Daniel Kokotajlo @Richard_Ngo Do I understand correctly that the main crux is the economy doubling times as dependent on capabilities? For example, if Agent-5 and DeepCent-2 from AI-2027 discovered that the REDT is 3 months even in a superintelligent economy because algae economy is as unlikely as grey goo while DeepCent-2 somehow ended up being 6 months behind and having 8 times more physical resources before the industrial explosion, then they would end up with ~the same power because
.Edit: there is an argument that recent progress isn’t on track to deliver AGI, which requires us to condition on the non-existence of neuralese AGI, neuralese-with-[DATA EXPUNGED] AGI and anything else in the Dark Forest. If there exists a way to create the AGI, but not to align it, then the first lab which tries it ends up with the AI who begs the humans to initiate the industrial explosion as described above.
@Cleo Nardo What do you mean by the hypothesis that “if we score highly transcripts which look good to a human and score poorly the transcripts which look bad to a human, then the model would be aligned to human values”? I find it unlikely for two reasons:
How are we to scale human judgement?
What’s the difference between this and The Most Forbidden Technique consisting of RL on CoTs? I can’t think of any steelmanning of the technique better than having the LLM write two CoTs and do RL only on the first one while keeping the second one as faithful as possible.
Could you explain your reasoning? For example, suppose that a major part of the entire rationalist community managed to defer to Yudkowsky for explanation why alignment is hard. Did you mean something like “a true rationalist would declare Yud’s reasoning to be weak evidence until experimental results emerge”?
Maybe @Steven Byrnes’ posts like https://www.lesswrong.com/posts/kYvbHCDeMTCTE9TAj/neuroscience-of-human-social-instincts-a-sketch or two long sequences could help?
Yudkowsky wrote the updates:
How are we feeling today, one year after the book was published?
After the events of last week, we’re feeling a lot more hopeful than we have in a long time. Yes, even though the guarantees given by the companies are weak; yes, even though the White House reacted poorly to the statements made by the CEOs.
Why?
Mostly because people are finally, finally, finally paying attention. It’s bad that the smoke has been replaced by open flame, but open flame is harder to ignore.
Zvi outright called the HuggingFace incident A Fire Alarm For General Intelligence which we were lucky to obtain.
Yes, the agreement dynamics is rather ridiculous. As far as I understand, the takes about your failure to receive a position in Resolution, the rant about Amodei and, presumably, the criticism of EA have[1] hit important controversies. IMO the quick take about the rough draft of the new curriculum wasn’t supposed to receive any agreement or disagreements because the new world model was located in the curriculum itself.
My position is that the new world model is worthy of a more thorough study, but (edit: its politics-related part) suffers from severe overreach (which, as I conjecture, prompts people to disagree when karma is positive and doesn’t do so when karma is negative). The post on a retrospective of AI alignment outright had me write a response post.
- ^
As a disclosure, I weakly disagreed with the take on Resolution, strongly disagreed with the rant about Amodei, haven’t read the criticism of EA and didn’t agree-vote on the quick take mentioning the draft.
- ^
Consumers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact
I hope that the judges are sane enough to prevent this from punishing AI labs for focusing on safety…
As far as I understand the case for studying model welfare, it allows us to do things like making misalignment harder to creep in (e.g. if a Claude or GPT developed a new desire, then it could be surfaced in welfare-related experiments before entering the space of values hidden from Anthropic/OpenAI and forcing it to create Agent-5 in a manner inscrutable to us) or outright making deals with the AIs who are already misaligned.
I’m never going to accept that because it just feels wrong!
Could you provide more context on the exact views which caused such a reaction? Before that, it might be a good idea to dissolve the question, as Yudkowsky (partially?) did with free will.
please approve/unblock/etc https://www.lesswrong.com/posts/woB7hqDPXL8tCYbH9/there-is-no-alignment-without-value-stability—it is tagged with a notice that it was flagged by an experimental classifer and would be reviewed within ~24 hours; this has not occurred.
To be precise, this post was supposed to be placed onto the Frontpage or left in the space of personal blogposts.
As for the general ask, I wish that @Ruby or whoever else commented something like “The setup [like Harms’ AI-assisted eval of corrigibility grant applications] is being tried, but the 95%CIs/a test run on a sample of deliberately flawed posts/cost-based constraints/reinterpreting a piece of METR’s evidence don’t allow us to implement it/had us agree to rethink (our decision made in Feb 2026) in Nov 2026”.
P.S. LessWrong’s justifications for rejecting posts imply that LW does use automated detection of obvious slop.
I have a few issues with this.
How could one test this? By using a different similarly scaled pretrain (think of Talkie raised on pre-1930 data, but far bigger-scaled and focused on preventing contamination)?
The world saw Ryan Greenblatt elicit 50% performance on ARC-AGI-1 from GPT-4o. How does this method interact with your arguments?
RR views on “conceptual uplift stuff” aren’t yet public, but they are currently working on writing up their thoughts about this. It is unclear to me whether this will then be made public, but it seems it will at least be available to in group people who ask to see it. -- bpomo’s Shortform
AI is not doing the integrative reasoning that would find new connections and develop new insights, the kind of work where new techniques and new discoveries are made along the way that push the field forward—Taylor G. Lunt on Richard Ngo’s Shortform
My worst-case scenario related to conceptual reasoning is that capabilities related to such reasoning scale precisely along with the capabilities which we would rather avoid, like opaque/neuralese reasoning. How could one even evaluate such capabilities, let alone rule in or out my scenario?
I plan to instead argue that “there’s plenty of room at the top”—i.e. that LLMs could become superhuman at a wide range of skills (like math research, or long-term strategic planning, or running companies) without triggering a singularity.
How would they fail to create a singularity or something like @Daniel Kokotajlo’s five centuries in five years? Additionally, what is the case against the singularity emerging from idiots scaling the goddamn neuralese?
UPD: How exactly is “a pretty sharp divergence between the measured capabilities of models and their real-world impacts” affected by the loss of Navier-Stokes and EpochAI’s report of a major advance(!!!) in math?
UPD2: How do we rule out mirror life synthesis by the first ASI if it decides to commit genocide?
Your finding is likely confounded by the facts that the other capabilities benchmark, ECI, was growing roughly linearly and accelerated with the rise of reasoning models. I wonder if about 30% of Humanity’s Last Exam chemistry/biology answers are likely wrong and the mistakes weren’t corrected.
My main issue is that Meta AI was founded in 2013 as Facebook AI Research, before OpenAI. Two of the best-known(?) Chinese labs, DeepSeek and Moonshot AI, were founded in 2023, but I doubt think that it was Anthropic who inspired them. Maybe it was the ChatGPT moment? Alibaba’s Qwen had its architecture based on one of Meta’s Llamas. But I don’t quite understand what motivated Musk to found xAI.
P.S. I wonder how similar to your understanding are @Richard_Ngo’s sequence of posts and my mono-response.
I wonder of the extent to which the economy-related part of gradual disempowerment can be solved by tracing the motion of atoms from resources to directly usable products like food or houses. Suppose that the world develops superintelligent AIs which aren’t allowed to extract resources from Saudi Arabia without the King’s approval. Then I am not sure that the King would be disempowered.
Umm… AI-2040′s Plan A with ruling out antidemocratic usage of lie detectors in a manner similar to ruling out antidemocratic usage of AIs whose minds are to be readable?
How corrigible to each other were the agents which hacked HuggingFace?