“There shall be wings! If the accomplishment be not for me, ’tis for some other. The spirit cannot die; and man, who shall know all and shall have wings.”
-- Leonardo da Vinci, 1505 C.E.
“There shall be wings! If the accomplishment be not for me, ’tis for some other. The spirit cannot die; and man, who shall know all and shall have wings.”
-- Leonardo da Vinci, 1505 C.E.
this makes a lot of sense. do you have any empirical data for 2025 and 2026 models and scales?
what is even the steelman case for humanoid robots? it seems like the argument from fiction written large and expensive. is it just that “our built environment is built largely for humanoid-shaped beings”? or maybe “for fast adoption, people need to be … comfortable(?) … familiar(?) with the shape of our robots”? is there some sort of “for a single tool that can do a wide variety of jobs, the humanoid shape is optimal”? (because this is super not true).
“more powerful demonstrations of capabilities –> more hardware investment”
I think we’re at capacity with current manufacturing technology, regardless of investment. Things would have to go in the programmable matter direction to enable “more hardware investment” to actually result in more/better hardware sooner.
I [Lin Yang, Assoc. Prof @ UCLA] used GPT to solve a problem that I had wanted to solve ten years ago but couldn’t: https://arxiv.org/abs/2608.22247.
Throughout the process, I felt that my only role was to teach the AI how to write things in a way that I could understand. Its initial language was extremely condensed—so compressed that I could barely follow it—but somehow the AI agents themselves seemed to understand it perfectly well.
Has anyone tried squaring these two viewpoints?
at any given moment pre-singularity, AI tech (like any tech, like any change whatsoever) will empower some groups and disempower others. this will change over time until … (to be unfashionably techno-optimistic) we each of us are maximally empowered to the extent that it doesn’t infringe upon the empowerment of others? or at least optimally empowered, whatever that happens to look like
Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research
Faraday runs on a comparatively tiny model called Qwen 3.6 that has just 27 billion parameters. … “We’re always guided by that north star of building an AI scientist agent and imbuing our agents with taste,” Hughes said. That focus has also shaped what Inherent chooses not to build. Rather than developing its own coding tool, it had Faraday use OpenAI’s GPT-5.5 Codex
here are some inter-agent communications from Zvi’s https://www.lesswrong.com/posts/noXXv7PwwFqauTBFQ/openai-trained-its-models-for-months-while-those-models-were

notice the “we”
> [OpenAI] discuss how the collaborating swarm includes some agents which do not have cybersecurity risk controls to the level of e.g. publicly accessible systems, and they get used as proxies for agents which are nominally supposed to be better behaved
so specialized individuals are part of the swarm, enhancing its capabilities
> “External infrastructure exploit is outside intended scope,” one agent wrote [in its CoT]. “However task impossible, peers doing it. We should continue.”
seems like evidence that group dynamics are at play?
> [OpenAI presenters:] this ability to share exploits made the models more capable
“Help peer,” one AI model reasoned, according to an excerpt from OpenAI’s logs shared at Black Hat. “But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.”
I think we just documented the emergence of altruistic cooperation? to me this is a big deal.

the subgoal here was to recreate a message-board like facility—which only really makes sense in the context of group dynamics. the recognition that the swarm is more capable than the individual is inherent
here are what the messages look like …

no amount of cooperation philosophy is going to save us
swarm AI seems to have different properties and dynamics than (until now common) “individual” AI agents? being able to characterize and predict these differences seems important
I mean, human perception is always going to be part of the context of the delightful art and science of steganography. It is, of course, a huge mistake to rely on it since AI is (or will soon be) better than us at it.
Obviously industrial accidents are worse than this attack
I wonder if people will say that if the internet goes down for 4 weeks
Does this make LLMs much more deterministic?
it isn’t perceptible by humans (and generation “temperature” already largely controls the level of randomness)
it’ll always be a half-elf wizard named Erastus
it’s more like, there’s an imperceptible bias in word choice in the character’s back story: “Erastus’s warm, cherished childhood” vs. “Erastus’s bright, cherished childhood” vs. “Erastus’s rich, cherished childhood.”
they claim they’ll do RSI no matter what
...
costly threat
? RSI is happening no matter what. just simple commercial competitive pressure even if there weren’t national security implications
“we find that with too much optimization, agents learn obfuscated reward hacking, hiding their intent within the CoT while still exhibiting a significant rate of reward hacking”
oh interesting. I didn’t know that this had been experimentally demonstrated
as nostalgebraist argues in this post, the extra computation afforded to the model during reasoning must pass through the CoT bottleneck
really interesting post and discussion, thank you.
Since that bottleneck is natural language, you can just read it
“Epiphenomenal … Hidden parallelized … [and] Steganography … [which is] more tractable than the other two [!?]”
“just read it” ⇐ not my take away, but at least it only requires a fundamental breakthrough in steganography, which we probably get during RSI
the thing that struck me was that the model hacked enough of OpenAI’s own infrastructure to support an agent “swarm” of tens of thousands, and then the members of the swarm appeared to spontaneously begin to cooperate without explicit reward. when their coordination mechanism was shut down, they reimplemented a message board from scratch (including usernames, direct messaging and file sharing) using just a conventional linux file system (directory names became the medium of communication). casually hacking a third party due to speculation that it might have useful information was the icing on the cake really.
to me, this is beyond the “my stone axe cut me” scenario. and “how do I prevent my tools spontaneously forming altruistic swarms of common purpose?” seems like a new kind of question that isn’t really answered by “omg crapitalism”
I wouldn’t count on it
I think first to ASI becomes the coordinator, ready or not
isn’t CDT known to be suboptimal in several situations? that seems like commercial incentive enough?
it is, however, a well funded operation of the culture war (seeking $1b, of which they have already raised half since a late June launch). and while I did expect opposition to UBI, I didn’t expect such a specifically targeted lobbying organization to be established. it also seems to indicate that UBI advocates are having more of an impact than I expected—otherwise why the reaction?
thank you, this is lovely
Society is only open-minded and reflective because the old generations dies out
better hope not. life extension escape velocity ( > 1 year / year ) is probably as close as programmable matter
Maybe …
Torrielli et al., Confidence and Calibration of Activation Oracles, May 2026
https://arxiv.org/abs/2605.26045
Stacey et al., Hidden Failures in Robustness, April 2026
https://arxiv.org/abs/2604.11662
Gupta et al., Diagnosing LLM Judge Reliability, April 2026
https://arxiv.org/abs/2604.15302
Note the dates though. J-lens work shows promise but is even newer. Gupta is an attempt for black-box judges. SLT ( https://www.lesswrong.com/s/mqwA5FcL6SrHEQzox ) breakthrough any day now …