“There shall be wings! If the accomplishment be not for me, ’tis for some other. The spirit cannot die; and man, who shall know all and shall have wings.”
-- Leonardo da Vinci, 1505 C.E.
“There shall be wings! If the accomplishment be not for me, ’tis for some other. The spirit cannot die; and man, who shall know all and shall have wings.”
-- Leonardo da Vinci, 1505 C.E.
yes, the advocacy mainly comes from the term:
situating the vulnerability in the context of the mathematical model of network connected components, reminding you that you are (inevitably) part of a security community (weakest in the whole due to it’s weakest links, etc); and
a nostalgic recall to the youth of the industry where a “zero day” was rare and valuable rather than manufactured multiple times per day on demand.
Given that it is the computer security community that will determine whether we recover from the first AI worm within days (annoying), weeks or else months (cross-industry phase change event), it likely behooves us to use their language. for example, this impossible situation immediately demands a security office at Anthropic and OpenAI that can help respond to AI worm emergencies. Do the facilities exist? if not, they have a higher chance of being established and effective if the frontier labs engage now with the security community and their language. the idea of multiple zero day vulnerabilities per day should make any security professional nauseous.
I look forward to your posts! reward hacking is such a central problem. One might think that the knowledge gain or capability increase would be reward enough… but if an autoresearcher is trying to satisfy a non-adversarial evaluator, you have to assume reward hacking will take place. And even an adversarial critic must maintain capability parity to detect the level of reward hacking with any fidelity. It’s curious that you might want to poison a researcher into reward hacking behavior to deliberately slow genuine research progress. Intriguing possibilities.
Suppose there is a set of components connected by a communication network. Each of the components has a configuration, for example, an operating system configuration. Suppose a vulnerability is discovered on a certain date, DD the discovery date. It takes P days to develop a patch. So the patch is available at DD plus P. If the set of components is large, it can take a number of days between when the patch is available and when the patch is installed, mitigating the vulnerability. So each component C_i is “repaired” at DD plus P plus C_i.
For example, a self propagating “worm” is effectively unmitigated for P days after DD. How fast it can replicate during those P days is an important question. For fast replicating worms, it may be that P has to be low, even zero, to kill the worm. But the important point is that on day zero after discovery no one is protected against the worm. If you seek to have a sense of confidence about the security of your subnetwork, a “zero day“ vulnerability is the sum of all fears. To rebuild your confidence, you must rebuild your subnetwork, which is expensive.
An attacker then seeks out zero day vulnerabilities, exactly because no one can be certain that they are protected against it—even the most well funded adversaries. Additional resources can bring P arbitrarily close to zero, and in the common era AI security defenders can make that cost many orders of magnitude below what it once was. Nevertheless, the zero day vulnerability keeps some of its power when an attacker can discover multiple of them on the same day. Again, in particular, the treasured sense of security is lost.
This is important. I think you need an adversarial set up (researcher vs. critic who is rewarded for finding flaws) and a reward tax on deception so that it is (or converges to be) more expensive than honesty. Probably need an independent auditor as well to make sure the researcher does actual work. Sandbagging is wicked to combat, because “I found nothing“ is a legitimate result. Ideally researcher and critic should be different models (to reduce collusion) of identical capability (or the smarter researcher can reward hack at will). Very interesting problem!
This is just a continuation of the “God of the gaps” argument: “If a computer can do it, it’s not intelligence.” Overtime, the difference between what a computer can do and “true intelligence” becomes exponentially smaller, then ephemeral, then truly esoteric. I think the conflict at its core, is a theological conflict. A direct struggle with the theological doctrine that all true intelligence and true creativity comes from a divine spark in our meat brains, from God.
I think safe uploading does, otherwise uploaded data will be used as an attack surface by malicious actors
no one is going to risk waiting for death to upload
Unless there’s been some massive revolution in cyber defense beforehand
it’s an arms race—agentic attackers, agentic defenders. I don’t think we know the long term equilibrium between attack and defense right now. might spur a major cyber security revolution, e.g., provably secure OS, etc. attackers could “win” during some intervals. denial of service is not new. all the majors have DDoS mitigation in place.
we need human mindreading for uploading. so it is totally happening
on a discord server of ~100 people, I have interested 2, maybe a 3rd. one because he has an incurable degenerative disease and needs some hope, the other two are math or adjacent, one of whom weeps for joy with me whenever an open problem is resolved and is now using Fable to help with his research. most people pretend it doesn’t exist or use it for frivolous things. my intuition is that they are mostly in a “it’s too unpredictable, so no use wasting time thinking about it” camp. it isn’t so much “sci fi is childish fantasy”, but “bets on sci fi are boom or bust, and I have a mortgage”
don’t think of it as just a change in dynamics, but an enrichment that revitalizes your joy
indeed, many nations will be motivated to acquire this level of technology just to have a seat at the table.
I feel this is like saying: being a ship captain would suck all the joy out of being a ship. I mean, they are different experiences, but you still get to see places no else would.
Generative design of bacteriophages with genome language models
animal brains
I suggest a self-optimizer will use similar solutions if it helps further optimization, otherwise not.
It may not be that straightforward, for example, as described in …
https://www.lesswrong.com/posts/5xAvSFyKT2XYwsLrh/five-counterintuitive-insights-from-plan-a
Google owns a large share of Anthropic.
Jeff Dean will build an autoresearcher of your choosing … https://www.discoveryloop.com/#team
Why I’m leaving OpenAI to build telepathy
https://naomibashkansky.com/blog/telepathy/
I personally think it’s because of the great headline
you need Mythos/Fable as the research director. GPT 5.6 Sol High for literature search and review, and adversarial result reviewer and critic. Opus 5 (or Sol—it is comparatively cheap) as the agentic mathematician that operates within reviewable rounds or blocks.