I endorse and operate by Crocker’s rules.
I have not signed any agreements whose existence I cannot mention.
I endorse and operate by Crocker’s rules.
I have not signed any agreements whose existence I cannot mention.
genes for ATP production in animals
Nitpick, but: ATP is not specific to animals. It’s “found in [and synthesized by] all forms of life”[1].
Although, interestingly, yesterday I heard about some bacterium that is so adapted to endocellular parasitism that it lost the capacity to produce ATP and relies on stealing it from its host cell.
Now I understood what you meant by “middling numbers”.
I haven’t thought about this in these terms, but yeah seems true.
I do think that expected values are generally valuable input into decision-making and sometimes good as the primary decision criterion.
The case that we’re talking about here one where one branch is not just low-value, but an irrecoverable loss (“doom”, extinction-level stuff), and that’s one of the places, where expected values tend to mislead the most. Cf. The St Petersburg Paradox, non-ergodicity (as Aprillion says), etc. The natural/obvious fix (unclear to me how principled/holistically good) is to apply expected value reasoning to the policy level, but it’s hard and I think it’s uncommon to do that.
(Not sure whether/how much you disagree, but I realized I had kinda assumed this to be obvious in the context, so thought it worth clarifying.)
In logical time, instrumental convergence precedes the terminal-ish goals that cause it?
Perhaps this is partly what you get if you promote reasoning in terms of expected values too hard, too generally (instead of “saner”, less extremistan-conducive decision rules, or even heuriatics whatever “normal people use”).
I don’t know how much this has to do being overly neoclassical-econ-brained, but it’s a live option for me that it does have non-trivial amount, whether it’s in terms of actually skewing people’s “honest thinking” or by giving them more concepts to produce palatable justifications for why they are riding the cancerous wave.
in active inference are beliefs which are fixed at an artificially high credence. We can then infer from those credences that we will probably take actions consistent with achieving those goals.
IIRC, it’s more like: there are probability blobs, which vary in their higher-level weight (“precision”?), and at the low-weight end you have paradigmatic beliefs, whereas at the high-weight end you have paradigmatic goals. Or something like that.
So I’ve updated my action from a 25% chance of training this month to a 75% chance. That does move me towards getting what I want! But there are two big problems here. The first is that (absent other constraints) if I want to win the race I should just train with certainty. The 0.75 number is an arbitrary artefact of my initial prior on whether I’d train; it’s not what a utility-maximizer would do.
The second problem is that if I fix the belief in winning the race too high, there won’t be any action that makes that credence consistent, and I’d need to change my other beliefs instead to restore consistency. More generally, setting some beliefs artificially high seems like it would propagate falsehoods throughout the rest of our belief web.
Going for the second one first, I don’t see the problem. P(win race next month) is contained to the [0.04, 0.36] interval by the setup of the problem (specifically by the conditional probabilities).
I don’t think the first “problem” is an actual problem either. P(win race next month) seems to be doing the work of steam, roughly, the amount of (available?) optimization power you’re applying to the goal. Abram actually writes that:
The simplest model of “steam” is that it is the agent’s own subjective credence that it will do something.
According to this, actions, policies, and plans will gain steam when you see evidence for them, and lose steam when you see evidence against.
You recover EU maximization if you fix P(win race next month) at the highest value allowed by the conditional probabilities: P(win race next month)=0.36 implies P(train this month)=1.
Also, the 0.75 number is not an artifact of your priors P(train this month)=0.25 either. The prior is irrelevant. If you start with the system of P(win race next month)=0.28 and the other two conditionals, then you recover the only consistent P(train this month) by solving P(win race next month) = α*P(win race next month | train this month) + (1-α)*P(win race next month | ¬train this month), and the only solution is α=0.75.
to me it seems like the natural response is to avoid fixing goals at all, and rather think in terms of “forces” that are trying to pull credences in goals upwards (which I’ll call drives).
As should be clear from what I just wrote, I disagree with your motivation, but I think it is a good direction for other reasons.
It sounds to me a bit like updating on virtual evidence from radical probabilism. Radical probabilism also rejects the rigidity criterion, i.e., that conditional probabilities must remain fixed across updating, which you also discuss right after.
It seems like belief webs implicitly implement EDT, which struggles to evaluate hypotheticals without interference from existing beliefs.
What do you mean by “interference from existing beliefs” here?
My longer-term hope is that belief webs will allow us to think of individual agents as an emergent phenomenon, rather than something we need to bake in when reasoning about intelligent agency. You could potentially consider all intelligent beings to be part of a huge, highly non-equilibrated belief web.
I kind of like this vision, but it probably makes sense to dispense with the term “belief” here. Active inference probably should have done this ages ago, and instead kept equivocating between [probability blobs serving function X and behaving in way W] and [probability blobs serving a very different function Y and behaving in way Z].[1] My understanding is that Friston similarly uses “preference” to mean something like “where the system ends up, by minimizing low energy”, which is a similarly flawed explication of the pre-theoretic concept of preference.
I don’t know whether “beliefs” you’re talking about here should be seen as probability blobs or something else. My guess, partly informed by davidad’s post, is that, on the one hand, pure probabilism seems clearly defective and overhyped in various ways, but probabilities also seem like “special building blocks”, in that beliefs are exactly functionals of probability distributions.
I am now less certain about my comment that sparked this post (especially about the “fire you after 1 or 2 months” part), but I remain rather unconvinced about the cluster of strategies you’re suggesting here.
I would update more strongly if I saw side-by-side-ish examples of something like what you’re proposing vs quitting loudly from harmful industries (tobacco? lead gasoline? social media clusterfuckery?). Admittedly, they might be difficult to find.
I think the best reasons to stay are (1) you can do good/beneficial work on the inside; and (2) you can try to influence the culture from the inside. (1) can work for some people, but it seems to me like they are — god bless them — an anomaly. I don’t expect (2) to reliably have a strong enough effect.
There’s also the issue of working in this sort of environment warping your epistemics a lot, for a combination of reasons, which gives you an additional reason to quit in order to reason more clearly about what would be good for you to do. Again, some people are fairly immune to this, but, again, I expect them to be an anomaly, and most people underestimate the extent to which social reality warps their cognition. @Richard_Ngo had a tweet about his personal experience of this, a few months/weeks after quitting OpenAI, saying something like, “I thought I could think clearly, even while I was working there, but oh man, I was wrong.”.
ETA: If someone thinks they have good reasons to think that they would achieve a better effect in this way, then I’m very happy for them to try.
ETA2: I’m also concerned that “just” having “safety people” on the inside acts as a form of safetywashing.
Assuming you mean “cheaper per unit of token/effort/something”, not cheaper in total to get the Lean code produced.
My understanding from the post, as well as my a priori expectation, is that formalization was a relatively small fraction of the total time/compute spent on this. They threw a fuckton of compute at the agent swarm and then wanted to save by switching to Astra for the final step? IDK, seems implausible.
I would also expect the new model not to be that much more expensive than Astra.
Bernie Sanders proposed his “ban ASI” bill a few hours before the release of Astra the Unmonitorable.
ControlAI proposed theirs 2 hours before OpenAI announced their Navier-Stokes solution.
Those guys have a great sense of timing.
Not to overstate it, but I find it a bit interesting that they had to use Astra, rather than this “more powerful” model, to formalize the proof once it was found.
Something something hard RL degrades skills that are not being RL’ed?
They would fire you after a month or two and the firing wouldn’t have the same social effect as voluntary quitting of, say, Daniel Kokotajlo or Richard Ngo.
Not confident about this, but I tend to think that this is less of an issue of something like “poor impulse control”[1] and more that environments like this one teem with the same sort of collective insanity as many other mundane human collectives: many political parties, many religious groups (not just “cults”), fraternities, sports fans, some friend groups etc. There you have stuff like:
perception of social reality immensely constraining what you can think;
people often have some awareness of that, but even those people often under-estimate the effect size
deference cascades;
inertia of collective action powered by a sort of identity;
some expected “pain” associated with the possibility of disinvestment from the company.[2]
On top of all that, you have evaporative cooling. Most of the sane ones have left.
Of course, I’m not saying that any of this is “fine”. But I expect that this will make it much harder to shift the public perception in the direction you’re outlining here, because it seems to me that it would be tantamount to awakening to the mundane insanity of so many human collectives.
On the other hand, insofar as most people typically don’t try to work on making their views on an issue coherent (unless they’re kind of exogenously forced to?), maybe it’s not as doomed? But what I’m describing seems like a track record/baseline reason for this to be very difficult.
I mean, you can, not-meaninglessly, generalize “poor impulse control” to being unable to properly cohere your short-term objectives with your long-term objectives or something like that. But, like, the point is that most humans can’t or don’t try to do that.
IDK how big a part of one’s social circle other lab employees are; something something status, including downstream from lucrative payment? I would guess that the biggest pain factor is some psychological sunk-costs-shaped barrier, but IDK
This seems to be aging very well and will most likely keep aging very well for at least a bit longer.
I was also thinking this, and that’s why I updated strongly in this direction upon first seeing this pattern in the tweets. But another option on the table is that they have a strong general policy of not leaking bits about internal R&D.
((Ooops, fixed))
The issue is the un(der)questioned even deeper atheism on the AI macro-strategic scale that implies that you should grab power because if you don’t, then the other less good guy does. As we all know, this has never ever backfired.
Dario rejects doomerism re:misalignment, fair enough.
But what of doomerism re:slowdown?
Furthermore, the last few years should make clear that the idea of stopping or even substantially slowing the technology is fundamentally untenable. The formula for building powerful AI systems is incredibly simple, so much so that it can almost be said to emerge spontaneously from the right combination of data and raw computation.
[...]
If all companies in democratic countries stopped or slowed development, by mutual agreement or regulatory decree, then authoritarian countries would simply keep going. Given the incredible economic and military value of the technology, together with the lack of any meaningful enforcement mechanism, I don’t see how we could possibly convince them to stop.Predicting the difficulty of cooperating around / enforcing a “substantial” slow down of AI development seems similarly difficult as predicting the difficulty of avoiding misalignment? Perhaps it is true that this would be historically unprecedented, but as Dario notes, the whole possibility of a country of geniuses in a datacenter is historically unprecedented.
[...]
FYI you linked his profile, not a specific tweet