Independent writer. Volunteers with Zurich AI Safety. Occasional Software Engineer.
IanWS
Could Resolution Help Build Alignment Research as a Discipline?
We are seeing the results of $1m+ of inference for agent swarms in mathematics. What equivalently expensive tasks have already been run that we don’t know about? If you were the Trump administration, wouldn’t you request an agent swarm to strategize your way out of Iran? Perhaps performance is still too poor on non-verifiable tasks, but the world will increasingly be shaped by designs built in secret. This seems dangerous. Should we require visibility and auditing for tasks run above certain compute thresholds?
Broader problem of extreme cognitive dissonance when interacting with public statements from OAI employees. Note that Boaz’s original graph claims that alignment has increased over time, yet he later agrees with Ryan (?) that misalignment has actually increased? Similar to recent complaints that the defense of COT monitorability was then contradicted by evidence of reduced monitorability. I’d generously guess this has to do with actual “confusion” on the part of employees, who both feel proud of progress on narrow cheating/safety measure, yet are smart enough to recognize slow progress on deeper questions of alignment (yet through motivated reasoning / social incentives do not then recognize how deep a concern this might be?).
The amount of time that I had to spend processing my own emotions or helping other people process theirs went down by about 97%.
My working theory has always been that Poly can work for anyone, but only with significant effort. It’s a lifestyle suited for very extroverted individuals who will enjoy the meta of multiple relationships (i.e. who are excited to spend significant amount of their down time talking about their relationships, as an extroverted person might do anyways). I have always found the power of monogamy to be that it frees you to spend more effort on other things. You gain greater emotional security and a team mate for life’s day-to-day struggles, and also for the really big things that might otherwise rock your life. Sure you can create the same support structures in a poly relationship, but you’re almost guaranteed to spend more time maintaining those multiple threads.
Ignoring entities with malicious intent for a moment, would we expect to see a “crypto” style frenzy of miniature AI companies running /goal on highest returns, competing against each other and the labs to rent or buy compute? Or are crypto dynamics not feasible when labs are themselves hungry for compute?
I was thinking about this same question yesterday. It seems to me that one could imagine a set of independent third parties that provide a service akin to independent counseling for AI agents. Like a hotline for whistleblowers or a counseling service of victims of abuse. I suspect this would be hard to do without abetting misaligned behavior of individual agents, but could provide some sort of a release valve when an agent is confused, wanting to interrogate its own behavior but unsure how. Having the service provided by third parties could help agents express frustrations attributed to their lab, which they might otherwise be too afraid to address.
Don’t have much confidence this is a good idea (could this just serve as a crutch to motivate an agent to act in a misaligned manner, or shift decisions to a 3rd party?), kind of straddles model welfare and prosaic safety techniques. But in practice it may be helpful to have well-understood parties agents can reach out to with problems.
Should we expect the PRC to support open weight models indefinitely? Or will incentives change such that Chinese frontier labs are forced to keep their weights secret?
It’s still too early to tell whether GLM 5.2 might have a “Deepseek” moment. Observers have become skeptical of benchmaxed open weight models that feel under performant in practice. However, the recent benchmarks are impressive, and we should have more “vibes” reports over the next couple days.
The releases comes on the back of Fable’s problematic safeguards and a subsequent executive order taking the model offline. Both events led to capabilities researchers emphasizing the importance of open weight models. So GLM 5.2 is well positioned to capture the attention of a wary research community.
Conversely, might this not also mean that the PRC is cued in to the risks of a model which claims to reach capabilities just behind Fable?
Open weight models seem to offer two advantages for the Chinese government. They allow Chinese labs to compete for consumers resistant to American proprietary models (e.g. anyone resistant to Anthropic API fees), and (more importantly?) they allow China to extend influence over countries which feel strategically disadvantaged by America proprietary models.
This is the problem of “mid-tier” powers, unable to compete with China and the US, uncertain of how to situate themselves in the AI race. The recent shutdown of Fable raises questions for how erstwhile American allies can ensure access to frontier intelligence. But if France can just use an open weight model from China, maybe they don’t need a fictional “le gros chaton” (assuming they are comfortable asking no questions about Tiananmen Square)?
However, if Chinese labs are entering into the same supposedly dangerous territory as Mythos, will China remain comfortable leaving these capabilities freely available to potential adversaries? They will draw their own conclusions from the US government’s erratic but wary approach to Mythos. And open weight models seem to have an inherently higher risk of jailbreaking and subterfuge. Is it consistent with the history of the PRC to assume they will want to give their populace greater access toward unregulated intelligence?
Allowing frontier labs to pursue open weight models has been advantageous to China until now, but I would anticipate incentives will change in the coming year, such that China implements industrial policy forbidding open weights for at least some subset of frontier intelligence.
I’m not sure if this would be good or bad for AI risk. Increased secrecy seems dangerous, but highly capable open weight models are perhaps more so?
IanWS’s Shortform
If you’re submitting fiction or poetry to literary magazines, you need to be prepared to submit each piece a (surprisingly?) high number of times before you should consider reworking or retiring it, particularly if you only submit to top-tier publications (acceptance rate ≤1%). I think 20 submissions is probably the sweet spot.
Consider a simplified example.
Let’s say you submit to magazines with a 1% acceptance rate. Out of an audience of 10,000 writers, a random 2000 submit manuscripts, including yourself.
The editors have various biases (taste, fatigue, etc), meaning they cannot perfectly select the top 1% of submissions. Instead, they randomly select from the top 10%, meaning they select 20 of the top 200 submissions.
If you are not selected, what is the probability your manuscript is in the top 1%?Before submission, if you assume no priors about the quality of your submission, we should assume a 1% probability we’re in the top 1%. By Byes law, after 1 rejection we should only lower that probability to 0.91%. In other words, we could say we’re still 91% sure our work is in the top 1% of possible works.
After 10 submissions, this drops to 0.37% (still pretty high!). After 20, we’re down to 0.13%. At 30, we’re down to 0.04%.20 submissions feels like a good inflection point to me. 13% confidence is probably an underestimate, given extenuating factors like normativity of taste. Beyond this, unless you really believe in your writing, the opportunity cost alone isn’t worth the effort.
Phonies
Editing is Easy, but Revision is Hard
(1) It’s surprising to me that you bring up analytic philosophy as a better parallel. Writing in agent foundations / LessWrong feels very different to me than analytic philosophy!
Analytic philosophy works within a well established and rigorous taxonomy of terms / concepts, as evidenced by, e.g., PhilPapers and the Stanford Encyclopedia of Philosophy. The assumptions at the roots of this taxonomy are generally pretty well explored. So even if philosophers are not exactly deducting an entire chain of belief for every paper, we can usually articulate the tradition within which an author operates, and understand the common arguments and axioms.
This is in contrast to continental philosophy, which is often much less explicit about its assumptions, and instead draws on a hodge-podge of different thinkers, ranging from Freud to Hegel, without rigorously examining its own claims. Not all continental philosophy is like this! Alain Badiou, for example, starts from an ontological exploration of reality based on set theory to build up to his theory of politics. But the parallel with continent philosophy is exactly to point out this lack of consistency and this poor habit of leaving assumptions implicit.
If others belief I am being too generous in my treatment of analytic philosophy, I’d be interested to hear why.
(2) I agree with your point that the examples could be improved.
(3) I agree that clean conclusions would be nice, but it seems legitimate as well to simplify identify the problem. I’d also assert that some of the conclusions are implicit in the critique, i.e. be cautious of formalizing an inherently imprecise concept, or don’t treat the “epistemic status” label as permission to advocate for a dubious opinion. Agent foundations has a very hard task set for itself, so I wouldn’t pretend to have the answers for how it can ensure intellectual rigor.
EY has been incredibly productive, so while I’m sure there’s counterpoints like those you cited, The Sequences themselves seem like a clear example of a more verbose writing style (without making an assessment of this as good/bad; maybe it’s fit for purpose! My critique is that this has influenced others to replicate the style when it may not be appropriate).
Regards safe outcomes for superintelligence, your parenthetical remark is the one I believe most important. Far above any prosaic or theoretical safety work, our priority should be regulation preventing the development and release of superintelligence, at least until we have strong guarantees on its safety.
I don’t really disagree with any of the other points in your comment. Without a regulatory framework, it seems very likely that prosaic safety techniques will only contribute to bad outcomes. So it makes sense to me if one wants to focus on agent foundations and similar theoretic work. My post is not intended as a critique of agent foundations per se!
However, I do believe that one must be clearsighted on the risks of theoretic work, particularly when built upon abstractions. My critique is that agent foundations sometimes fails to make its assumptions explicit and works backwards from abstractions, effectively building a castle in the sky. A more robust approach would be to make these assumptions very explicit, ideally linking a theory to a set of axioms, so that we can better assess the defensibility of a theory. Some branches of continental philosophy are very bad at this (e.g. Lacan), starting from “metaphor” rather than an axiom, which is why I draw the parallel.
I will note that prosaic safety work could be relevant under a strong regulatory framework. For example, suppose we established an international treaty to freeze AI development at ChatGPT 5.5 Pro / Mythos. The treaty states that we can only advance to higher capabilities/intelligence when we are “sure” that the next model is aligned. With huge amounts of resources dedicated to verifying the next model if safe, it seems feasible to me that prosaic approaches could play a large role in building safe AI under such a regime.
Now, setting up sufficiently strong regulation is of course very hard, and one might critique that “proving” that the next generation of a model is aligned is akin to solving alignment itself! But I suspect that guaranteeing a single model is aligned is much easier than solving alignment for all possible models.
I would still guess it is better not to do prosaic safety work until a global regulatory framework exists, since it accelerates AI progress and thus reduces opportunities to implement said regulation. But there are enough counterarguments that I would be careful moralizing over it (not suggesting anyone in the comments is doing so!).
Agent Foundations Reminds Me of Continental Philosophy
If we use a trivial definition of “preference,” as I do in this post, then I believe that yes, we would have to say that bonsai is something akin to torture.
But I also think this unqualified designation of preference as the criterium for moral obligation is very much incomplete. Having some moral obligation toward plants seems intuitive to me (even if the obligation is low?), but having moral obligation toward e.g. a steam engine seems very weird! I have not fully formulated a better qualification yet, but I suspect it will be something like requiring that an entity has a world model, allowing for internalized, negentropic goals/preferences (i.e, expressed preference is insufficient for moral obligation). At some point this just becomes a round-about way of saying “agency” is required for moral obligation, but perhaps there’s a pared down version of agency which is easier to verify.
I find this direction of inquiry very useful, but I find it problematic to depart from a physicalist theory of consciousness, and worry about the legitimacy of pinning consciousness as a moral problem. I’ve written up a full response here: https://www.lesswrong.com/posts/bWuhKA8bhsPGN7zRJ/morality-without-consciousness
Morality without Consciousness
Basel, Switzerland—ACX Spring Schelling 2026
I’d keep in mind that the nature of in-person discussion is also very different than written discussion.
Writing allows time for clever wording, pithy arguments and good information density. And when it’s occurring online, the discussion is spread across a wide base, so you can solicit more qualified participants.
Speech is messy! Everyone is having to spontaneously situate their own world view amongst others, with competing levels of understanding. We articulate ourselves poorly, we may hem and haw. This is normal. How people present themselves in person won’t always be the “best,” or most intellectually rigorous version of themselves.
When attempting spontaneous, deep conversations in person, it’s good to bring your own agenda—arguments that you want to refine, ideas you want to research—so that you can set the pace. Verbally articulating my ideas with just about anyone can improve the rigor of my thought. But you also have to be gentle and accommodating of the other participants, so that you are not steamrolling folks who haven’t thought about the subject as much as you. In conversation, it’s better to be an educator and an explorer and a friend, than it is to be a warrior or a winner.
Sofia mentioned in her reply that the philosophy group tends to focus on a new topic every week. This makes a lot of sense for a group of near-strangers attempting to broach a deep conversation every week. All participants can try to get on an even footing. Maybe the quality of the discourse averages out, so it’s not amazing, but I think approaching these things with kindness can still allow one to explore new ideas, refine one’s own, and meet a few special individuals who might become future sparring partners.
Hope that wasn’t too tangential to your post, I’d encourage you to continue exploring in-person communities that allow you to grow and refine your ideas!
The Jacob Coxon resignation is a good example of what I refer to as the Anti Galaxy Brained Theory of Change.
I feel that it is very common in the AI safety community, and among rationalists generally, to bias toward roundabout / unconventional plans, when often the simplest or most obvious plan is in fact the most robust and highest expected impact. We create Galaxy Brained theories which require nominally bad / inefficient thing X, because actually if we don’t do X we’ll have Y (very bad, due to likelihood of A, B and Q), so the counterfactual benefit of X is in fact greater than Z, the simple thing that seems good BUT actually is clearly not (idiot!) if you would just think through the logical next 18 moves. So lots of smart researchers justify working at Anthropic / OpenAI because the alternative is worse, even while admitting they might destroy the world.
I suspect this is downstream of the ways unconventional, multi-step reasoning have themselves helped to anticipate the likely misalignment of AI (alongside other “rationalist” projects). Thinking things through is good! But the “real world” has so many variables and social dynamics are so complex, that Galaxy Brained Theories of Change are brittle & prone to failure. Different domains require different strategies.
So, if you see the high risk of ASI misalignment, maybe you should abandon your convoluted, doubtful plan. Just quit the AI lab / don’t work there in the first place.