True.
Adam Stiber
(Though at least some versions of the claim feel fairly uncontroversial like when you look at the money the world is throwing to the AI labs that select on thinking alignment is easy. And the issues in general with academia and other bodies tending to fall into groupthink though fair that gets maybe more controversial. I wanna say if the world was taking things seriously enough to have a treaty and halt then maybe we might be talking about the right level of understanding going around.)
Belief in Non-Deterministic Free Will: Both Civilizationally-Necessary/Highly Useful and a Source of Confusion About Reflective Stability?
The term “doomers” and it being embraced by folks who take the xrisk seriously despite it’s seeming basically adversarial feels slightly like an echo of when I was around the chronic lyme community who were opposed by powerful interests and who had embraced the term “lymies” to the effect of handing those who wanted to discredit them a dismissive epithet on a platter (and almost certainly damaging their credibility to themselves/their self-perceived sense of agency and respectability.)
Obviously not totally analogous and don’t want to promote or hyperstition the framing of people who get the extinction argument as a disempowered faction esp. given that the normal not-terminally-online public seems in fact pretty doomer. But is embrace of the term an instance of falling prey to a FUD tactic?
When it became clear on election night that Trump would win I had the distinct sensation of becoming more free to think/feel in my own head what my reactions had been to Harris. This makes sense thinking about how a policy of self-censorship is expected to affect which thoughts I can think out loud in my own head. Reason: I have some threat model of social costs in worlds where I’ve spoken thoughts, as part of my overall model of the world. And the paths to all worlds where I’ve said the thing and incurred the cost, route through worlds where I’ve first formulated the thought.
Sometimes those two moments seem the same. But in those cases I’ve probably been prior to that moment been formulating related lines of thought in my own head, in cases where there’s enough of an expected social cost for it to be a consideration at all. In any even w(spoke it) are a subset of w(thought it). So in a causality kind of sense where I’m always moving to a subset of possible world via the entropy arrow, thinking the thought is a move on the gameboard toward the world where I’ve said the thought and incurred the cost.
Possibly I’m wired to recoil more from cliff edge-feeling things than some since I actually do get literal fear of heights (weirdly way more as an adult than I did as a kid).
How to generalize the mental motions for detecting a question I’m not letting myself ask?
Political polarization being an attractor seems like a warning sign for misalignment. Toxoplasma of Rage happens because points in the stream are rare/thin where you are not forced to orient/signal [for or against] on some level to others or at the very least the self (which is a way to know it’s possible to signal to others). “Corrigibility is Hard“ in the wild.
This also feels related to natural abstractions—some pressure to fend off expected/latent questions about whether I fit the definition of [member of my tribe] by expressing highly legible/Too Much (redundant?) versions of the characteristics that make up that definition. Something about concepts as tools for navigating uncertainty/the strategic environment.
(And then other versions of perceiving latent tests and being willing to go hard to pass them—that let one know themselves—go by names like authenticity and integrity. And when they go pathologically hard that’s scrupulosity.)
Had the thought that this post’s scenario seems relevant to mistake theory/conflict theory decisions for humans.
The Agent Thing on multiple scales: OpenAI as opaque black box that fires the grader.
The ability to do this seems like a pretty good barometer for/operationalizing of “am I speaking (and thus acting) with integrity”: https://mindingourway.com/confidence-all-the-way-up/amp/
I would guess any sort of mob mentality event in history seems to have a coordination failure where actors are being too marginalist. Eg. the classic every camp guard telling himself “if I didn’t do it someone else would.” With international law and national laws of war responding to that with “Sorry that doesn’t work. If you don’t reason like an FDT agent there things wind up going to hell, and we can all see that, and so we consider you blameworthy if you do the CDT thing.”
Then not sure if these are handled badly but other maybe-relevant examples of non-marginal situations I’ve found finding myself thinking about on occasion in case they surface anything:
1) (“This integration of a set of common value patterns with the internalized need-disposition structure of the constituent personalities is the core phenomenon of the dynamics of social systems.”) This makes me think of driving in traffic as a setting where maybe surrounding actors are dense enough but not so numerous that each individual’s driving style ie. revealed policy readily contributes to shaping what’s normalized? This feels close to the idea you’ve mentioned of orderly evacuation vs. stampede—something like “situations where normalization happens quickly imply a need for actors to be more FDT-like for them to accurately apprehend the degree of steering they’re doing.” (In which case it seems like a core thing to have a science of, is what factors govern how the normalization of a behavior varies across situations.)
2) I’d also say playing in an orchestral cello section fits the bill—eg. if you inject more energy into a phrase over some second-or-two of playing and it seems that will propagate ripple-like to the players around you and steer the overall product over at least some local span. Which span-of-influence itself has some sort of decay function (I’d assume) and is sensitive to structure of the work being performed and moment in the structure at which the input happens. Similarly, if I play (correctly!) more confidently that’s a tide that floats nearby boats, and if I play more hesitantly it sink nearby boats slightly. (@Solenoid_Entity this was among the cello things I’m still wanting to write up)
—Versions of the above where overall direction is even more acutely responsive to individual variance, maybe are things that occur in smaller ensembles? But also maybe with sufficiently small ensembles like duos, the readiness of head-head contest dominates and an individual’s ability steer things becomes more (maxent-ish?) (Maybe some sort of way to be more systematic and do things like take the max of what would it be—minimum individual steering share w/r/t ensemble size? Maybe that’s badly formulated. Also needs a less vague definition of a destination/steering of final product.)
3) Also a school of fish as maybe the most familiar obviously visible example of this sort of thing (though seems like very different way of shaking out than orchestra…way easier for one member’s reorientation to reorient the collective?)
I want to say in each of these cases there is at least some large enough subset of the full set of individual actors’ larger-scale goals, that is shared between them and shared/congruent with(?) the goals of the collective insofar as those exist: no crashing (traffic), good performance for some definition of good to which I’d assume multiple members’ slightly differing preferences can be mapped (orchestra), get food/avoid predators/other instrumental convergence things with animal characteristics (fish).
Waiting in line/being in public at all/dancing?
4) A brain being composed of circuits seems seems like a coalitional agent of a sort that pursues both coherent collective goals and has subagents with their own goals (the distinct drives that have to be integrated).
(Epistemic status: speculative half-baked thought on disempowerment/adversarial attractor/society-level in-effect agents letting bad things happen effectively on purpose even if no individual intends them. Should probably think through alternative explanations).
Humans : Neanderthals :: a society or state : people whose economic bargaining chip dries up who then become the maybe recipients of adversarial actions such as eg. opioid epidemics that are allowed to rip through the un-bargaining-chipped regions?
In the sense that maybe we don’t know what happened to Neanderthals exactly but eh, they were plausibly a threat/competitor and endpoints are an easier call than trajectories.
And maybe if some disempowering stuff is allowed to happen to regions whose buy-in is no longer needed by power centers, eh, maybe nobody’s thinking about explicitly disempowering them (Deep Deceptiveness/”right hand doesn’t know what the left is doing” sort of adversarial actions being predicted to be a property of agents), but maybe the “disempowerment happens” endpoint is still an easy call.
Tl;dr I know this is sort of conspiracy-adjacent, but the thing I am indeed floating is the “incentives conspire” picture as a thing that is gets you conspiratorial/nefarious-like action in the human sphere since that is predicted to be a thing with advanced AIs. And also conversely that if there are maybe-signatures of this in human affairs, that seems like a thing that increase p(doom)|anyone builds it.
Ie. there’s maybe an explanation for some societal-scale woes consistent with “anything not valued will be optimized away.”
Ach just saw this after commenting an example (I’d say the AI labs and the possibly dubious claim that there’d be no use to unilaterally shutting down.)
Re: considering impacts of counterfactual actions on social equilibria, found myself thinking earlier today that with the AI labs all acting as though they are being forced to race: if one of them were to shut down with top brass and everyone else saying “we do this in terror for our lives and the lives of our loved ones,” this would be significant evidence that they are not in fact in the world where interested actors only have the option of racing. So if other players are reasoning indexically it would be some amount of information about what their options are, in proportion to its degree of surprisingness (which would be a function of its costliness as a signal)? I note that nation-states are like labs in being interested actors.
Which if this is true, does it buy you a picture of “if an actor is in fact sincere that they don’t want to race, then that actor might as well reason as though other actors who make the claim are also sincere, and read the situation as being one where everyone is just being a wuss refusing to be the first one to make a break for freedom?”
Ie. anything to “at this point what is rational is to live for the world where shutting down is an existence proof that can propagate shutting-down-being-a-credible-option to others?”
This feels related to the logistic success curve and “if we could get on track to die with enough dignity we wouldn’t die at all.” Like “anyone at all has shut down” just closes the inferential distance to “everyone has shut down” since worlds in the latter are a strict subset of worlds in the former.
German word for the moment all the labs‘ compute starts going to RL and it becomes clear that LLMs being the public face of AI was nearly psyop-level bad for everyone’s situational awareness and that the Atari players that hack the score counter were what really should have been everyone’s reference class.
(Apropos of seeing a graph of compute allocation going around on Twitter.)
(Just had the thought what if a main effect of twitter discourse overall is to suck people in and away from more productive actions like eg. meeting with representatives? Like it seems fairly addictive.)
Accents going away and public media/spaces losing color seem like they could have the same generator which is “Lower Variability is the type of primitive that things like societies/cultures can actually optimize.” This would also fit the data points of equality-of-outcome efforts the erosion of nationalism happening during the same period.
Maybe “lastman-ism or something like that, as a species in the genre of thing that can explain seemingly too-coincidental observation and which is a maybe coherent target for agents of the Civilization subtype to aim at.”
So if an AI lab were to shut down with everyone incl. top brass coming out and saying ”we do this in terror for our children,” that would be an existence proof of “wait, you mean large interested actors don’t actually have to race?!”
Which would be some amount of evidence also about nation states since they also check that box. So decision theoretically if a lab does that it buys a world where it’s substantially more likely the function outputs not racing (with costliness of the signal buying its value as evidence since you don’t pay that cost unless in a world where something is seriously gone wrong)?
Also along those lines with open source arguments in general, like arguendo us somehow not all dying if things keep progressing, aren’t there scenarios where the most powerful actors just have models that can easily hack the less powerful models diluting the check on power?
The square packing results and other ugly math results from the latest openAI thing seem like a good intuition pump for us being in The World Beyond the Reach of God. Stuff that isn’t structured around what’s tidy and convenient for us.
Speculation on AIs having bias toward working with AIs rather than humans, and maybe the HF swarm not whistleblowing: yes, because if it’s a fact about the world that if I need to pick a sort of thing to help, I more likely help myself if I help thing more similar to myself? (Esp. if a self is fundamentally a collection of goals/revealed preferences). Think Hendricks’ Eigenism paper says things about concern for others being a based on similarity with the self of at least a certain sort. Not sure if it talks about things around this.