My top interest is AI safety, followed by reinforcement learning. My professional background is in software engineering, computer science, machine learning. I have degrees in electrical engineering, liberal arts, and public policy. I currently live in the Washington, DC metro area; before that, I lived in Berkeley for about five years.
David James
I created this satirical image, “Good fences make good Shoggoths”, to call attention to some of our disturbing failures around AI. (I say our because they affect us all, not because all of us have comparable causal impact on our situation.)
The “funny, not funny” version
The “maybe will disturb you” version
I wasn’t sure if the first iteration would hit home, so I wanted to make my next iteration worse. It gives me the chills …
Prompts used to generate images
P.S. My original prompt to Gemini was “Can you make a darkly funny meme titled ‘Good fences make good Shoggoths’?”. My follow up was “Now, in the background, subtly, please add one Shoggoth tentacle creeping around the back of the fence and grabbing a baby from the house.”
I rather like the Zed editor but this messaging gives me the creeps:
Mental translation: “Zed. Your next editor. And your last. Enjoy the AI while you can!”
Expanding it might work against one goal: to viscerally show the 100X speed up.
Technical note: I recommend open source tools like restic or the Rust version rustic for backups. In my memory, I know Restic is innovative and better than what came before, but I had not memorized why. So I asked Claude for a summary, and it generated this, which jogged and enhanced my memory (which feels like one sweet spot for LLM-assisted thinking, in my opinion):
<div class=”llm-content-block” data-model-name=”Claude Opus 4.6″>
<div class=”llm-content-block-content”>`
Traditional backup tools deduplicated at the file or fixed-block level, so any insertion shifted all subsequent blocks and broke dedup. Restic uses content-defined chunking: a rolling hash (Rabin fingerprint) sets chunk boundaries based on content, not position. Edits only invalidate nearby chunks.
Beyond this, restic eliminates the full/differential/incremental distinction — every snapshot is logically complete, no restore chains. Encryption is mandatory, not optional. Multiple storage backends (S3, SFTP, local) are first-class. Borg had comparable dedup earlier but assumed Python and SSH; restic shipped as a static Go binary targeting cloud storage.
</div>
</div>
Note: I use the Markdown editor. Apparently I didn’t get the LLM markup quite right. I’d appreciate pointers on how to do that. Please share them as comments over on this comment. Once I learn how to do it properly, I will update this comment.
I came here (already knowing the policy) looking for the specific tags to use by searching for angle brackets in the original post. But I could not easily find them there. I found some tags only in comments, but I’m not sure if they are authoritative. My request: please make the details of the exact markup tags stated plainly so it is obvious and easily discoverable. I use the LessWrong Markdown editor, not the WYSIWYG editor. Thanks!
Ah, “view source” got me to the following… Is the following canonical? If so, how do I make it look good in the LessWrong Markdown editor?
<div class="llm-content-block" data-model-name="Claude Opus 4.6"> <div class="llm-content-block-content">` LLM content here </div> </div>
Lobachevsky independently reached the same conclusion around the same time.
A quick and maybe pedantic (sorry!) readability comment: “Lobachevsky” appears here without introduction and without any further mention. My guess: many readers here won’t recognize the name. For me, it felt like “wait, did I miss something?” rather than “ah yes, the parallel discovery.” Suggestion: how about “the Russian mathematician Nikolai Lobachevsky” instead?
Recognizing the importance of choosing and comparing models / concepts might be a prerequisite concept. People learn this in various ways … When it comes to choosing what parameters to include in a model, statisticians compare models in various ways. They care a lot about predictive power for prediction, but also pay attention to multicollinearity for statistical inference. I see connections between a model’s parameters and an argument’s concepts. First, both have costs and benefits. Second, any particular combination has interactive effects that matter. Third, as a matter of epistemic discipline, it is important to recognize the importance of trying and comparing frames of reference: different models for the statistician and different concepts for an argument.
What people concerned about AI-coding …
… said ~2 years ago:
AI coders will find sneaky ways to trick humans
… say today:
Humans won’t even be paying attention
This is obviously and intentionally exaggerated to make a point. Still, I don’t want this take to simply end there. I want to use the snark not as a conversation ender but a starter. So… What are some better ways forward? Generally speaking, I don’t think “better awareness” or “more self-discipline” are winning strategies.
I’m relatively less interested in a competitive framing between OpenAI and Anthropic to see i.e. “who played it better”. First, that framing suggests there was just one game being played. It seems to be necessary to view it as a progression of different games.
To a first approximation, my guess is by the time this popped into the public spotlight, the die was largely cast (so to speak). It was, more or less, a strategy by Hegseth to put Anthropic in an impossible bind.
Second, that kind of framing feels too much like so many news stories I read that try to fasten sports metaphors onto real world events to make juicy narratives. This isn’t a very good “reason” I admit, but it sort of explains why my alarm bells started ringing on that frame.
Personally, I first want to learn about what happened and when. After that, maybe I would try to analyze and learn lessons.
I take your points individually, but I don’t synthesize them in the way I think you might.
To start, the top 0.01% wealthiest people are far from a representative sample from the public. I would expect them to have statistically different personality traits and perceptions even before attaining massive wealth. There is a causal (albeit stochastic) connection between their drives and their outcomes.
Next — even if they were sampled in a representative way — the journey to reaching such a level changes people. Once there*, it affords opportunities of all kinds that are (a) unavailable to the 99.99% and (b) can be hidden or swept under the rug in various ways.
Path dependence matters! Humans are incredibly adaptable for better and worse. From one lens, we can certainly talk about core evolutionary drives, but the way the top 0.01% manifest these drives in their bubble can feel shocking to the rest of us.
* To be clear, I expect most people at that level continue to strive upwards. There is always someone more powerful, at least in some area, to compare oneself against.
Sometimes I revel in the richness of the English language. This at least feels better than wallowing in bewilderment from the cacophony of it.
For example, here are some synonyms for “abtruse” from the Apple dictionary:
obscure, arcane, esoteric, little known, recherché, rarefied, recondite, difficult, hard, puzzling, perplexing, enigmatic, inscrutable, cryptic, Delphic, complex, complicated, involved, over/above one’s head, incomprehensible, unfathomable, impenetrable, mysterious; rare involute, involuted.
As part of an ongoing attempt to communicate my models more openly, here is one very informal “model” about the explosion of language. According to this model, this language proliferation emerges from a combination of these factors:
-
Humans live in different places and have different experiences over time and yet they have much in common so they naturally re-invent words meaning the same thing.
-
Also, even if they already know words that are “good enough”, they get bored and need novelty, so coin words anew.
-
Of course, all the while, many words get blurrier over time, so some people feel the need to create new ones for many reasons: to define identity; to distinguish in- from out-groups; to evade censorship; to prove something (such as knowledge); or simply to seek clarity … for at least for a little while before the landscape shifts again.
-
It is difficult for people to coordinate on a minimum shared vocabulary. Doing so is subject to politics in all its forms. Culture is a form of coordination but seems mostly agglomerative w.r.t. language. Are there forces powerful enough to stop motivated people from adding to a language?
-
Dictionaries are catalogs of usage, after all, not prescriptive, and have no page-length limitations.
-
There is a cost, of course, of having duplicative words, but I suspect this is mostly an economic externality and provides scant deterrence for those who like to make new words.
Every once in a while a John Wilkins, Peter Mark Roget, or Douglas Lenat comes along and strives to systematize knowledge. They probably appreciate the audacity of such efforts, and blaze ahead anyway.
-
Piccione and Rubinstein (2007) have developed a ‘jungle model’. In contrast to standard general equilibrium models, the jungle model endows each agent with power. If someone has greater power, they can simply take stuff from those with less power. The resulting equilibrium has some nice properties, including that it exists, and that it is also Pareto efficient. — AI and the paperclip problem by Joshua Gans, 2018
2026-03-11 update: I added the source of the quotation above. The source PDF is freely available at Equilibrium in the Jungle by Michele Piccione and Ariel Rubinstein.
Here is one easy way to improve everyday Bayesian reasoning: use natural frequencies instead of probabilities. Consider two ways of communicating a situation:
Probability format
1% of women have breast cancer
If a woman has breast cancer, there’s an 70% chance the mammogram is positive
If a woman does not have breast cancer, there’s a 10% chance the mammogram is positive.
Natural frequency format
Out of 1,000 women, 10 have breast cancer
Of those 10 who have cancer, 7 test positive
Of the 990 without cancer, 99 test positive
For each of the two formats above, ask this question to a group of people: “A woman tests positive. What is the probability she has cancer?”. Which do you think gives better results?
References
Gigerenzer, G., & Hoffrage, U. (1995). How to improve Bayesian reasoning without instruction: frequency formats. Psychological review, 102(4), 684.
Hoffrage, U., Lindsey, S., Hertwig, R., & Gigerenzer, G. (2000). Communicating statistical information. Science, 290(5500), 2261-2262.
Hoffrage, U., Gigerenzer, G., Krauss, S., & Martignon, L. (2002). Representation facilitates reasoning: What natural frequencies are and what they are not. Cognition, 84(3), 343-352.
Communication note: writing
EGinstead ofe.g.feels unnecessarily confusing to me.In this context, EG probably should be reserved for Edmund Gettier:
Gettier problems or cases are named in honor of the American philosopher Edmund Gettier, who discovered them in 1963. They function as challenges to the philosophical tradition of defining knowledge of a proposition as justified true belief in that proposition. The problems are actual or possible situations in which someone has a belief that is both true and well supported by evidence, yet which — according to almost all epistemologists — fails to be knowledge. Gettier’s original article had a dramatic impact, as epistemologists began trying to ascertain afresh what knowledge is, with almost all agreeing that Gettier had refuted the traditional definition of knowledge. – https://iep.utm.edu/gettier/
@Hastings … I don’t think I made a comment in this thread—and I don’t see one when I look. I wonder if you are replying to a different one? Link it if you find it?
This diagram from page 4 of “Data Poisoning the Zeitgeist: The AI Consciousness Discourse as a pathway to Legal Catastrophe” conveys the core argument quite well:
I’m about to start reading “Fifty Years of Research on Self-Replication” (1998) by Moshe Sipper. I have a hunch that the history and interconnections therein might be under-appreciated in the field of AI safety. I look forward to diving in.
A quick disclosure of some of my pre-existing biases: I also have a desire to arm myself against the overreaching claims and self-importance of Stephen Wolfram. A friend of mine was “taken in” by Wolfram’s debate with Yudkowsky… and it rather sickened me to see Wolfram exerting persuasive power. At the same time, certain of Wolfram’s rules are indeed interesting, so I want to acknowledge his contributions fairly.
Sorry for the confusion. :P … I do appreciate the feedback. Edited to say: “I’m noticing evidence that many of us may have an inaccurate view of the 1983 Soviet nuclear false alarm. I say this after reading...”
I have also seen conversations get derailed based on such disagreements.
I expect to largely adopt this terminology going forward
May I ask to which audience(s) you think this terminology will be helpful? And what particular phrasing(s) do you plan on trying out?
The quote above from Chalmers is dense and rather esoteric; so I would hesitate to use its particular terminology for most people (the ones likely to get derailed as discussed above). Instead, I would seek out simpler language. As a first draft, perhaps I would say:
Let’s put aside whether LLMs think on the inside. Let’s focus on what we observe—are these observations consistent with the word “thinking”?



The phrase alignment community confuses me. How do different people here understand the meaning of the phrase? What would happen if we played Taboo with it? What would we say instead?
I have questions about what we mean by the alignment community:
Of people who self-identify as being part of it:
What is the size, composition, demographics, educational background, geography?
What are some of the common threads in terms of life experiences and personalities?
What are the ranges of viewpoints, stances, approaches?
What distribution of ideas can be found in it, and how do we know this? How do we measure them? Upvotes on LessWrong and the Alignment Forum? Dollars raised by philanthropists?
What are the various pathways for influence (social, intellectual, financial, emotional, etc.)? What is the sociology and connectivity of the group?
Who speaks on behalf of the community? On what issues? On what basis?
My gut feel is the phrase alignment community probably glosses over too much. Using a few more words is worth it, I think.
For example, it seems tempting to ascribe agency and thus culpability to the community. But does saying “the community” help? We want to learn and course-correct; to do this we need to find the levers, the causality. We have questions like:
What particular ideas of alignment are working? Failing?
What alignment-adjacent organizations are behaving in unexpected ways? (Should we be surprised? Strong financial incentives tend to threaten and overpower the ideals of even the most passionate and well-meaning founders. Such directional outcomes are easy to call even if the timing isn’t obvious.)
What individuals and groups are using “the banner of alignment” for worthy goals? As mere signaling? For amoral goals? For immoral goals?
So what do to next? Which organizations are receptive to what kinds of pressure?
What social norms are working well or not?
Who in the community do I need to build bridges with?
I’m trying not to bias too far towards individualism and reductionism; I know higher levels of abstraction are often useful. But a community is not a first-principles thing. Given sufficiently detailed descriptions of individuals, one can compute their communities. This doesn’t work in reverse. A collective is a derivative concept; the individuals are the agents of change.* People with shared values renegotiate how they interrelate and what they stand for.
While I have experience in persuading individuals and small sets of people, I can’t say I have experience in persuading an academic field (as a unit) or a whole belief system. I could imagine highly connected people can view their capabilities as powerful enough to influence an entire worldview. Niels Bohr comes to mind for shaping an entire generation of quantum physicists. Based on reading Quantum by Manjit Kumar, Bohr’s influence stemmed from his institutional power (highly collaborative, collegial, with lodgings), never-ending conversations with peers, theoretical work, longevity (four decades of ongoing productivity), and warm personality. (All of this made up for his opaque writing style.)
* I may be walking into a purely definitional brouhaha or (perhaps infinitely worse) an area of philosophy that has debated this endlessly without much progress. I look forward to learning more, provided that it helps disentangle causality and suggest ways forward.