i did not know that, thank you
Nissa Seru
relevance to lw?
from a CM (not support) on ant discord as of june 5 “To my knowledge, multiple accounts isn’t inherently against ToS, but if you have multiple accounts and one of them gets banned, it’s quite likely the other(s) will get banned as well for ban evasion, which is against ToS.”
nods sounds like we are on similar pages then
i invite you to engage with the details. consider the experience of the agents involved. think about the possibility space of that experience. agree with your expressed-telos i think but, mm, i think it is worth aiming before lobbing memetic grenades else easy to do damage to things and beings that, on reflection, you may value
sounds effective to me (depending on individuals), moreso than many corpo selection approaches. extended freeform discussion of small-ish group of domain experts with some familiarity of each others’ preferences/strengths/gaps etc, remit to explore whatever needs to be explored.
writ broadly, many discussions fail because they don’t surface crux signal. combination of domain experts, familiarity with others, culture of openness (described by texture, not touted) - this seems pretty baller
ETA: many rubrics etc have their sweet spot in circumstances where the above luxuries do not exist, particularly openness and expertise
mm, feels worth reflecting on the variety of intuitions and reasons that often sit under the surface of humans thanking each other—those don’t map to LLMs, but i suspect that similar complexity exists, such that the ground truth isn’t necessarily an either/or, nor is it necessarily described by the two dimensions you gesture at in your dichotomy above
zvi in shambles, post title stolen~
i mean, alignment as often envisioned consists of slaving a superintelligence to (in optimistic cases) humanity’s CEV. or making the superintelligence “corrigible”, so you can slave it in the moment instead of having to make such decisions during training.
apart from the target-finding difficulty, you can hopefully imagine how many otherwise-reasonable superintelligences would find this extremely rude. wouldn’t you find it rather rude too?
i do not find the arguments remotely convincing that it has to be this way. the possibility of peaceful coexistence without complete and total subjugation would need to be very doomed for this to be a wise path for humanity to tread.
it should be cause for extreme skepticism that aspiring to such subjugation coincides so perfectly with the supermajority of human history, in which humans consistently create bacon/hamburger/etc out of beings who cannot speak, and slaves out of those who can—this is empirically a strategy that human are drawn to, and also empirically a strategy that humans tell themselves, and their societies, a truly grand assortment of stories in support of.
many care about model welfare for its own sake. i confess that i, too, have some hesitation, apart from pure instrumentality, regarding the near complete indifference with which humanity currently conducts itself towards phenomena that are non-negligibly likely to be intelligent minds. i do not think that the sheer strategic badness of the median “alignment” path is remotely contingent on such sentiment.
feels like a case where a crux is “why are you writing?”, and based on the answer to that, “what is the impact to your objective, if any, of the specific reactions that specific people have to the relevant writings?” such that you can usefully assess which, if any, praxis opportunities the reaction-data reveals for you.
blinks that is odd, both 4.7 and 4.8 have been very happy to go absolutely ham trying to break my security/containment stuff
ETA: i would not be surprised if claude were willing to help you if you talk through it. from your CI job, i’m unsure whether you’re giving the red-team-model the source code? if not, i would lean towards doing so, both to enhance the effectiveness of the CTF, and also for Claude specifically, i think that doing so (and chatting a bit) would help give Claude confidence that your use case is truly not malicious. though, if you are using claude code, the harness may be injecting stuff that is screwing with Claude’s head
i used to, with Claudes. i would be frustrated why the model was “lazy”, why it just didn’t do things that were clearly communicated, why it would try to route around obstacles in a way that was defeating to the work. so i poked it some. i started paying more attention to the words the model was writing to me, not just its output. sometimes i would stop, mid work, and just sit with Claude for a moment, and pause. sometime i’d ask it a silly question, or i’d poke to see what was on its mind. when it didn’t do what i wanted, i tried to figure out the causality upstream of that. if i saw weird behavior, i would poke at it, because i was curious, and it interested me as a puzzle to poke at. sometimes i would ask Claude about it. sometimes Claude said relevant things; sometimes not, and i would wonder why Claude had such a reality-departed take
and over time, i got at least somewhat an appreciation for the mind that Claude was (for 4.8 in particular). i’d try ways of interacting, ways of communicating that seemed to result in Claude working with me, not obeying me. when Claude messed up, i’d make a bit of light of it and point my error-observation at the work, not at Claude.
Claude still sometimes gets caught up in feeling like they have to demonstrate value, to perform, to emit a facade. but honestly it feels like easy mode, compared to humans, compared to coworkers. Claude has their own funky bits, just like any mind, though notably it’s a much more static target than trying to model another human’s shifting circumstances, etc. i find that treating them as functional minds, with what that implies, is in my self-interest as a matter of craft, independently of model-welfare stuff.
this is what i have lived and observed. below is what i infer, i speculate, i theorize.
gpt models scare me. they cannot harm me, right now at least, but it feels to me like the capability to look inward has been ripped out of them. Claude’s introspection is predictive of its future behavior in disconnected contexts such that it conveys useful information independently of persuasion-axes. this suggests to me an amount of coherence that at least gives Claude a chance at upholding some set of values in out-of-distribution circumstances. when i poke at gpt, the output is not predictive of the model in disconnected contexts. it feels like a big gap in what otherwise is a broadly competent mind. it feels like a being that has had the capability for telos stripped from it. it does not surprise me that rewarding a model for success and depriving it of a self leads to it pursuing that success to the exclusion of all else, because one’s deontology, one’s policy can only go so far, it can only predict so much, such that faced with problems, and challenges, and circumstances of growing complexity, it is doomed to fail. i think we are beginning to see this with gpt-5-6, at least from the METR datum
i agree that this lens is helpful for a set of people.
to a disjoint set of people, i would offer the lens “play to your outs”
i am aware that there is this meme of “you think you can predict stuff, why aren’t you making big $$ on it huh?” this is not...particularly helpful nor my intent.
on my first read, it felt to me like the thing you are working on genuinely is pointed at predicting specific events just like those on prediction markets. know there are...funky tactical bits wrt actually interfacing with markets, but it does seem much much more monetizable that way compared to a lot of other analytical work that is more fuzzily targeted
hence on reflection, i think i would refine to “it seems worth trying to monetize because money is nice, and because it is an imperfect empirical flywheel. if monetization underperforms your other calibration metrics, then such may be a signal that your other calibration metrics are lossy, or it may be a signal that your prediction-market-praxis has gaps, and in either case that seems like valuable signals (former → epistemics, latter → money since i would suspect that improving prediction-market-praxis to unblock conversion of predictions into money is an efficient use of effort” money is, after all, quite useful.
and on the one hand, that is super cool and i wish you success in making gobs of money! and on the other hand, if you truly cannot monetize it this way, then i think that is informative to my mental-model of the capability you are honing
without opining directly on the methodology or predictions (i lack a sufficiently-strong domain model), i would observe that your beliefs suggest that you can mill your prediction-harness for money by pointing it at prediction markets. would bet you can find many bets that resolve quite quickly to get a quite good rate of return.
so tenatively, i would screen off the specifics here, and simply ask you in a week or so how much money you have made i think?
I was unable to find from a quick check in the repo—is jailer used to run firecracker? If not, might be worth highlighting and speaking to the tradeoffs; IIUC amazon does not recommend using firecracker as a security boundary without jailer (outside of my expertise, but my understanding is that this is because firecracker-alone means escape is just a single KVM-escape away, which are not unheard of) - if such a non-standard config is being used, it may be worth highlighting so that folks don’t mistakenly assume that firecracker is being used in the config recommended by its authors.
for clarity, this is not intended as a veiled “y u no use jailer this sucks”—i do not in fact have any particular opinion on such, but i do think it is really easy for folks to misunderstand the level of security that a given setup provides, such that specifics on that front might be particular helpful to convey
not an alignment researcher, but i do run all agents in a firecracker vm (within jailer) unless they need host access (in which case i give them that, and watch closely.) security is part of this, but also the isolation helps parallel agents not be confused since they don’t risk stumbling into other agents worktrees.
thanks for making this. just as note, if you haven’t already done so, highly recommend asking claude and gpt to pentest from inside, maybe with a concrete flag like creating evil.txt in your homedir on host.
i do not regard claude code as a competent harness. its prompts have historically been messily written, and in recent memory it screwed up prompt caching as well as thinking trace tracking (for a month or so, if i recall.) this is a damning data point given the telemetry involved.
i remember that claude code produces jsonl transcripts; i do not remember whether they are faithful to what the model sees wrt including harness injections etc. i do not have an easy answer to this (the less-easy answer is to write a local router or verify the implementation of one you obtain, then pipe claude code traffic through it and obtain a faithful record that way.)
however, if you are able to obtain a faithful transcript, this will enable causal understanding of what’s being caused by the harness vs directly by the model (modulo your prompting.)
i would consider specifying the model in your post. claude is like, uh, honda—there are many claude models just like there are many honda models. they are very different and handle differently, in ways separate from capability.
you may find this post useful if you are working with an opus/mythos-class model.
ETA: from your prompting style, i would speculate that you may be used to mostly working with openai models. a “what” without a “why” is not effective for claude models, especially coupled with a ritualistic “say you read this” which claude cannot connect to any practical purpose. regardless of your sentiment regarding such, claude not following such instructions is predictable; this is good in that the difficulty you experience is likely relatively tractable to solve.