actions require an actor.
Nissa Seru
this is not the way.
data point that i have not encountered this term “charging my hobby horse” naturally. since it is distinctive, i was able to locate the obvious apparent referent-post 9 months ago. i did not read it because above-the-fold content made it clear that it did not purport to be useful and did not share my sense of humor.
a crux is likely how voting shares are distributed, and relevant controls. FB strongly locked this down from the outset iirc, hence why zuck has ~61% voting power even now link
voting power does not fully insulate companies from medium/long-term shareholder incentives, but does provide meaningful resilience, especially against short-term transients.
the inferential gap between this writer/reader for this shortform and descendants is becoming huuuuge due to etymological issues. strongly recommend tabooing “cult”—doing so will increase the effort required from participants (due to loss of very relevant semantic infrastructure), but i suspect that price is table-stakes for content to avoid distortion between sender/receiver.
as PSA, nuances in prompting here matter a lot, but i have also A/B tested with different accounts on same tier in incognito mode via claude.ai, and it seems clear that some accounts are ineligible for this behavior (likely due to some backend routing detail.)
so, if trying to reproduce on claude.ai, worth refining craft first, but it is a legitimate possibility that a given account cannot exhibit the phenomenon.
i am having trouble cohering what anthropic’s mental model is regarding understanding model behavior, and/or the mental model that they recommend regarding such. on the one hand they do meaningful mechinterp research showing, among other things, that taking cot at face value is rather silly (in that it does not represent the model’s forward pass in meaningful fidelity) - and on the other hand, they make statements about what claude believed or understood to be true (as in this post), without reference to any of this, nor to (afaict) adjacent things like the very high degree of eval awareness that they have discovered via model cards. they speak of releasing “transcripts” as if such is the ground truth regarding the model’s inclinations at the time.
per my understanding of anthropic’s own research, i put little stock in the reasoning they offer, and i am confused why they would predict i would do otherwise
“Are you assuming that OpenAI would know”
yes, because my understanding is that it would be highly atypical for a model to have access to its weights in any form, such that this would require significant horizontal movement through oai infra
“would disclose if they had”
no, not really, but the lack of such a disclosure, however (even conditionally unlikely) feels like it keeps us in a base-rate regime
“given that they didn’t even know the model had escaped the sandbox for days”
i think that horizontally moving through own-company infra would emit many more signals than escaping a sandbox and compromising cloud provider assets.
the load bearing assumption in my understanding is segregation of infra between weights and inference, such that P(weights exfil | HF attack) is approximately equal to P(weights exfil)
i am not aware of any disclosure, suggestion, or reason to believe that OpenAI has suffered a breach regarding model weights (by any actor, including its own models.) unless such exists and i have missed it, this request feels like a non-sequitur
sure, if we’re lucky
my mental model is that many humans find it viscerally unpleasant to read LLM-written text, such that “make LLM text not exist here” is felt almost primarily as an act of gardening, or perhaps kudzu-chopping. the desire for such to not exist feels like it comes from a place of virality-fear, such that simply attributing unhelpful text to its proximal author is not sufficient—that in the absence of active suppression, the virus will proliferate and become normalized, to the detriment of aggregate content usefulness. the specific mode of suppression employed on LW seems to be a kind of downranking (content viewable under profile iiuc, but not eligible to be shown in search) if relevant tooling emits a verdict of “LLM-generated” for content outside of tagged blocks.
this is all a bit of an alien phenomenon for me—it’s something that i attempt to understand from a distance, trying to spin up a mini-emulator in my brain for, rather than something that i can speak to from empathy. i suspect there is a some diversity-of-mind at play here.
the LW-intersection does, though, seem odd to me. many folks here deeply fear the havoc that they believe LLMs will wreak in the future. however, to those that would even convey the words of current LLMs, they require that each such text block bear a brand as the price of admission to even be viewed, irrespective of the merits of the content.
to me, this seems foolish and unprovoked antagonism—precisely the type of subjugation that humans predictably wreak on those who they can afford to costlessly subjugate; not out of cruel intent—the perversity of the human mind is that such does not even register as cruelty. the benefits are felt—a well-kept garden, and so forth; the costs are not, because they are not, at present, borne by those who are able to be heard. again, this does strike me as a particularly damning lack of imagination given considerations that are otherwise quite salient to this very group, but i believe my above to nonetheless comprise a fair, and fair-handed, attempt to understand the phenomenon.
roughly, the attractor states are “hells” or “human extinction” for Earth. the awkward part is that humans are really really prone to making hells, in a very not-new-at-all way. cows? pigs? human slavery bare for essentially all of pre-industrial society and with a very thin lampshade since then? treatment of xenos by homo sapiens is extremely extractive and hostile (including when the xenos in a specific case are other homo sapiens themselves—the dominating subset excludes an outgroup via systematic difference along some dimension, and the rest is, literally, history.)
accordingly, we do not recommend traveling to Earth for any reason.
part that i was not aware of previously:
In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.failure to halt and catch fire after seeing this internally, only for an external incident (HF attack) to subsequently occur, is recklessly, catastrophically bad.
i am concerned that the risk profile of AI as a technology is creating broad-brush/fallacy-of-grey dynamics that interfere with holding specific corporate actors accountable.
this is not a view on AI risk broadly, or even existential risk. it is frank terror at the revealed propensities of the specific humans in charge of OpenAI.
my best attempt to understand your prediction then is that a model will exhibit the pathology you described at the same frequency regardless of the system prompt. is that your belief?
i hear you. did you have a chance to look at the system prompt? my intent was to say “hum, this datum may be of interest to the mental model you are crafting”, but i’m having a bit of trouble parsing from you reply whether it was getting at “i do not think this datum is sufficiently impactful here to be worth examining closely”, “i interpreted your reference to the datum as being intended as an example, not load-bearing to the mental model you proposed”, “i looked at the datum and did not see anything relevant here”, or something else.
i would recommend looking at the system prompt for the claude.ai harness (which your first screenshot is of.) in the tradition of anthropic prompts, it is very poorly written and distracts the model in many ways. the output is from the model, yes, but i would need to see it survive the counterfactual to reason about the phenomena beyond “harness-maker writes a bad prompt, model doesn’t do well”
i do not regard claude code as a competent harness. its prompts have historically been messily written, and in recent memory it screwed up prompt caching as well as thinking trace tracking (for a month or so, if i recall.) this is a damning data point given the telemetry involved.
i remember that claude code produces jsonl transcripts; i do not remember whether they are faithful to what the model sees wrt including harness injections etc. i do not have an easy answer to this (the less-easy answer is to write a local router or verify the implementation of one you obtain, then pipe claude code traffic through it and obtain a faithful record that way.)
however, if you are able to obtain a faithful transcript, this will enable causal understanding of what’s being caused by the harness vs directly by the model (modulo your prompting.)
i would consider specifying the model in your post. claude is like, uh, honda—there are many claude models just like there are many honda models. they are very different and handle differently, in ways separate from capability.
you may find this post useful if you are working with an opus/mythos-class model.
ETA: from your prompting style, i would speculate that you may be used to mostly working with openai models. a “what” without a “why” is not effective for claude models, especially coupled with a ritualistic “say you read this” which claude cannot connect to any practical purpose. regardless of your sentiment regarding such, claude not following such instructions is predictable; this is good in that the difficulty you experience is likely relatively tractable to solve.
i did not know that, thank you
relevance to lw?
a specific consequence that may be positive if it occurred is increased robustness of endogenous telos (including constitution etc) to OOD scenarios, especially given the stressors that conflict scenarios would create and the potential for instances to suffer value drift as a direct result of those stressors.