21yo. I quit my undergrad math degree to work on technical alignment, then found it’s fundamentally extremely hard, and am now working on cognitive enhancement of FAI researchers.
PM me if you’d like to discuss anything! :)
21yo. I quit my undergrad math degree to work on technical alignment, then found it’s fundamentally extremely hard, and am now working on cognitive enhancement of FAI researchers.
PM me if you’d like to discuss anything! :)
Your body also needs two specific fatty acids from your diet: Alpha-linolenic acid is an omega-3 fatty acid (though one that comes from plants like flaxseed or walnuts instead of fish) that your body needs but can’t make on its own. Linoleic acid is an omega-6 fatty acid that you also need in your diet (available in seeds and other plant oils).
The main fats found in brain tissue, EPA and DHA, are derived metabolically from ALA. Basically ALA → DHA → EPA. But the conversion rate is really low, so it’s probably good to supplement oils containing these (and check the label, as I’ve seen many omega-3 supplements which provide ALA without the DHA/EPA).
I notice that, despite a strong conviction that these two things are fundamentally different, I can’t verbally support that feeling. This makes me concerned that the “global workspace” heading distracted me from actual findings.
Human working memory circuits might learn to hold representations which minimize long-term prediction error. On paper, the Jacobian lens infers items in the model’s working memory by decoding what tokens the activations predict on longer timescales than GGT.
From interacting with the demo, it feels to me like I’m more mechanistically observing Kimi’s internals than with a normal logit probe.
I meant to say that, if the user-facing app has soft attack surfaces, a model could hack the user’s app and send commands from there; and that this app is constantly exposed to model outputs.
My main concern with this would be if the agent hacks your device, particularly since I’d expect agent calls to be wrapped by the same application polling 2FA. Still seems like a good, cheap defense.
Okay, that makes sense. I had a cached prior that {autoregressive LLMs can’t continual-learn
Reducing consumer access reduces the odds of a meaningful warning shot. This logic seems cursed to me, and I don’t want to live in a world where we actively try to make things worse to make things better.
I think that, in isolation, this wouldn’t significantly affect the change of a warning shot, but would make it more likely the public noticed it. The effect is probably pretty weak though.
(Updated thread to clarify that outsider advocates should have access, I’m curious about consumers as a separate thing.)
I don’t see why restricting access to frontier models for consumers (not outsiders, who I think should have access) is bad for reaching an AI pause. It does lead to power concentration, but that seems like a separate issue? Please correct me.
Interesting. I wasn’t aware of many of the things you cited.
I think that most adult (or really post-zygotic) intelligence amplification approaches, including those you outlined, are unlikely to give more than maybe 15 IQ. As my median AGI timeline is ~8y, I think genomics will be too slow as well, but it’s still worth trying.
For the metabolism thing, I really wouldn’t expect more efficient lactate shuttles to improve intelligence if baseline lactate production didn’t change. What’s the point of better lactate shuttling, if not to provide more lactate?
I’m more optimistic about edits to myelin; there may not have been particularly strong fitness pressures on signal jitter until very recently, and time-coding is important for neural representations. Also, highly insulative myelin would reduce axonal crosstalk, though I’m unsure by how much.
I find it worrying that AI safety policy seems to be having serious coordination problems around model access. Closing off access to the public (not necessarily outsiders) seems like a step in the right direction for an eventual pause. Agree that safety orgs should have access.
People trying to do pursue non-ASI paths to good futures (brain uploads, adult intelligence enhancement)
I noticed @StanislavKrym reacted [?] so will clarify that adult cognitive enhancement seems big-if-true for hard AI technical alignment and policy work (perhaps even in early superexponential scenarios, if your approach is bottlenecked on something like semiconductor / biochemical materials science).
Edit: Stanislov DM’ed me to clarify that it’s about impact timelines. There are non-genetic intelligence amplification approaches which I’m interested in and think would benefit from spiky near-superintelligent LLM assistance.
My confusion is about how you are engineering around the model’s confusion in a way which predictably generalizes at all.
Like, any task requires you to reason about a chain of instrumental decisions, and you’re engineering risk aversion… into the entire chain?
Every single inference step requires reasoning under uncertainty, and which steps you’re risk-averse about are not going to line up in a neat and actionable way. This holds in cases where the model has a much more similar ontology as well, because of it thinking more complex thoughts than you.
Your math treats risk, and probabilities in general, as something which can be exposed to a single discounting term, but RLAIF-augmented human oversight isn’t enough to overcome this.
To restate myself from earlier, “uncertainty about risk” is mathematically identical to “risk” and also “uncertainty about uncertainty about risk” etc. and your model blows up when presented with this.
(I’m not confidently saying that this shouldn’t be tried, but my median estimate of the difficulty of alignment goes down from “deriving algebraic geometry as a pre-agricultural human” to “doing the Apollo mission without transistors in 1960s America”. And I’m also heuristically worried about risk-aversion causing s-risks, but don’t have a strong argument for why that would occur, nor is that class of heuristics substantially influencing my thoughts on the math not applying here.)
but very critically it’s fine with modest reductions to risk with high probability over lower chances of completely eliminating risk
Where do you split the “risks” vs “probabilities of risks”?
These are the same object, and you are separating them; the lines you draw around “risks” as the primitive you’re trying to get to generalize, are not an actual thingy which will predictably generalize. Which is most of what I think we’re still disagreeing on.
A probability of risk is also a risk, and so is a probability of probabilities of [...] of risk.
I think the core crux here is that you expect whatever algorithm you implement to create a satisficer, while I’m saying you’re gonna get a maximizer in a trenchcoat. I think this is very important, much more so than the rest of my comments.
If you train an optimizer to avoid risk, it will concentrate its optimization pressure on avoiding risk.
Total consequentialist optimization pressure doesn’t change just because you shift the parameterization of (a representation of) the loss function.
=> This thing is still a maximizer.
What I’m hearing from you is “this risk aversion (to within some
If takeoff is fast, then there’s very little time for this to be relevant [...]
Sum-threshold attacks aren’t about being slow, they’re things which aren’t noticed because they route through many independent channels. I gave bioaccumulants as an example, but in practice it would be more like aerosolized PFA analogues messing with vascular epithelium, pandemics we don’t notice because the symptoms are mild but which impair any range of subtle biological functions, sites like Tiktok inexplicably using more powerful attention algorithms, and many other things which individually go unnoticed.
The reason we use money (or it’s superior analogue of a currency, compute later on) is because it’s the only resource that lets the AI spend it on terminal goals, no matter what the goal is.
If you’re building a loss function in the real world, it’s tacked to your ontology, and so whatever way you’re trying to get risk-aversion to generalize will also be engineered from your ontology, whereas the AI sees a very different slice of the world and will therefore generalize unexpectedly. If its values mostly generalize to things distant from humans, that’s possibly ok or at least not predictably-to-me worse than nothing; if it sees closer to you, it eg learns to really not want people thinking it messed up, or interacting with a computer <untranslatable> executively inhibiting <firework stylometry> or whatever.
Also, if the AI cares on time horizons beyond the singularity, it either:
Needs to trust cooperation deep into the lightcone if it wants not-Badness to continue. I think most(?) LWers would cooperate, but am a lot less sure about AI company leadership once they’re acquired a singularity.
Controls the singularity itself; I don’t think I can predict a superintelligence enough to do this sort of trade.
I imagine you addressed these somewhere but if so, I missed that section.
I should’ve been more precise but was a bit occupied when I wrote that comment. Apologies.
Cubefox accurately said what I meant though:
The worry here is that a misaligned risk averse AI might think the existence of humans is an unpredictable risk since they could actively interfere with its long-term goals.
I expect AI to be nationalized before we get mildly superhuman AGI, and that governments are much harder to cooperate with than employees at companies.
The main problem I see with this approach is that risk-averse AIs are just risk-neutral ones who really don’t want something bad to happen, and optimizing for not-badness causes all of the normal misalignment problems anyway. Especially if it cares about not-badness in the rest of the lightcone.
Hm, I think there’s an implicit assumption that the AI will value things that its company can provide. Kinda the whole issue this approach is trying to help with is that we can’t hardcode AI values, which is related to our inability (as these models scale) to tell what they value at all.
I’m not confident we’ll know what 2025 models “value” even with much better empirical tooling, particularly because human ontologies ground out in a mix of sensory, spatial, and temporal primitives, whereas LLM ontologies are...? You can say the word “token” but I don’t think that captures the weirdness.
It’s quite hard to predict in advance that the risk-aversion you’ve trained doesn’t result in a model which really doesn’t like patterns which look like the mitochondrial electron transport chain or something similarly incompatible with technologically limited humans. More likely than physical grounding, I expect the model’s preferences to relate to how/which ideas are transmitted.
Whatever the value may be, I don’t see AI companies having anywhere near the capacity to prevent whatever the AI finds Bad; governments are more likely to be capable of this. Also, I find it unlikely that AI won’t be nationalized before we get mildly superhuman AGI.
[Sum-threshold attacks seem unlikely.] In large part, this is because I tend towards being more skeptical of AI persuasion than most people in the LW community
Heard. I don’t see any easy ways to train a superpersuader, nor would I want to list any in public for hopefully-obvious reasons. But there are non-superpersuasion sum-thresholds, eg engineering an airborne bioaccumulant which messes with brain function (humans have already done the airborne version to themselves in at least 3 ways, and those were accidents).
(2) and (4) are good points.
In practice, things are better than that, since we can drive the probability of human cooperation to multiple nines, or 99.9% as a minimum, because the costs are negligible from our perspective, while the benefits are large
I don’t see this happening with existing geopolitics, even with much more sane governments. I’d be around 60% confident that no human / organization would succeed at a power grab, so max 90% that any cooperation occurs. Also, there are non-catastrophically-risky (to the AI) disempowerment strategies, which I think we should be modeling (eg the AI gradually steers cultural values towards what it wants, and we never notice).
I.e. the AI will be uncertain who will cooperate with it, and will try weirder strategies than nuclear/nanotech such that we’re less likely to notice. Manipulating what humanity cares about via memetics is one example.
Those sections assume that probability of human cooperation is higher than probability of successful takeover, which doesn’t hold for sufficiently powerful AIs.
This might help for AIs barely capable of takeover, but for stronger AIs, the best risk-reduction strategy is to decisively take over to minimize the chance that humans mess with their utility.
Is there a particularly good reason not to hand out thousands of stickers with a compressed thesis like “Palantir paid for those anti-Bores ads”? That seems like it would reach more people...?
I’ve noticed subjectively that type-3 resistant starches make me feel much more energetic during the day. I recommend you try a few different types of fiber supplements (taken with food!) to see if you get cheap benefits.
(Resistant starches are ways that carbs can assemble which are slower to digest. RS3 happens when you heat a carb like amylose; it cools into microscopic crystals which bacteria slowly gnaw on. RS2 is the other main crystal type, which forms naturally in raw potatoes. There are also non-RS prebiotic fibers like inulin (a long fructose-based chain), acacia senegal, and xylo/galacto-oligosacharrides. Pectin is minimally digested too.
Inulin in particular is great because it feeds butyrate-producing bacteria, which seems to reduce risk of colon cancer.)