MIRI, formerly MATS, sometimes Palisade; views mine
For more information, please reread.
100 percent AI generated according to pangram.
Yes; IABIED has been frozen at #669 on Amazon all day which, as I understand, means your sales velocity has approached implausibility and Amazon has frozen the number to check for fraud.
[Ordinarily this number changes hourly.]
Can you reply to this comment with another screen capture of the chart once it starts to dip again?
Sales of IABIED are also surging significantly.
Did you click the ‘view methods’ buttons next to the various graphs? That’s where the most interesting information lived in my read through.
His previous public post does not say “[I] joined OpenAI primarily because [I] like working on timelines”, which is what your (very flat) comment says above.
That would be the kind of public attestation that makes this a reasonable move (e.g. Matthew Barnett talking about his personal lust for immortality and his valuing of existing human lives over future lives, making it fair game to chalk his decision to work at Mechanize up to that source).
As is, this is just your supposition, building on a sliver of his public statements while ignoring a significant chunk of them, as well as a significant chunk of his work (as he points out in his response to you above).
For the avoidance of doubt: I too am critical of Thomas’s decision here and attempted, at length, to elicit his reasoning in a semi-private setting. I have not, broadly, had much regard for his strategic thinking over the past several years (while I acknowledge, of course, his technical talent and breadth of knowledge, and personally enjoyed talking to him the handful of times I have).
Notice the difference between Thomas’s response to your comment and his response to Jeremy’s: a thumbs up react. If you think your comment is equivalent to Jeremy’s, please consider why Thomas would react so differently to yours.
Your relationship with Thomas is not Jeremy’s relationship with Thomas. Notice that this moment was Jeremy respectfully building on a pre-existing series of private discussions directly with Thomas, and was not the first time Thomas heard ‘Jeremy believes x about Thomas’. Note also that Jeremy surfaced multiple points, rather than flattening to ‘you’re just excited to work on y’, and offered some pointers to the shape of his other disagreements.
Your engagement and relationship are both dissimilar to the example you cited; afaict this is a total dis-analogy. I understand that that might not be legible, and why it may have misled you about the norms here.
I deeply object to this kind of psychoanalysis of your direct interlocutor as a conversational tactic, especially in a public forum, and do not consider it an endorsed norm on lesswrong.
If you have the kind of relationship with Thomas that would lead to you having privileged evidence of his specific psychological composition or motivations, then you also have the kind of relationship that empowers you to bring those concerns to him in private.
A norm of viewing lab employees as epistemically compromised is fine, but guessing publicly about the exact shape of a specific private individual’s motivations is a profound violation, absent direct public attestation of those motivations on their part or substantial public evidence and load-bearing cause for the collation of that evidence (eg someone shifting from private to public life, or being caught in specific direct lies).
wild speculation: maybe it’s a version with adjustments to its maximum serial depth (via tweaking the recurrence hyperparameter), which then receives additional post-training so that it can make better use of that higher recursion factor (similar to post-training models to make better use of a higher CoT budget than they were allotted at other stages of training).
(this is a special case of your chart 2)
Thank you for saying this! I’m pretty sad that this is the state of things, and think it’s pretty likely that this basic disagreement (specific threat models vs respecting Goodharting and unknown-unknowns), is also at the root of many other contentious issues / past wounds on either side.
I’m glad that you’ve made these updates, but I also hope they’re allowed to propagate further (both within you and without you). There’s a very wide range of issues on which I’ve ‘~taken Eliezer’s side’ (with, regrettably, varying degrees of tact), downstream of this exact divide.
I think the deeper actualization of even a small move along this continuum would likely significantly change the actions of most people in the Constellation intellectual lineage.
The agents may not naturally connect natural latents to Condensation, a different take on a similar idea, with some exciting recent breakthroughs (I don’t know the natural latents literature well enough to know what fraction of Sam’s Condensation work is represented there).
The new name doesn’t do the job. I think a less-defensive version of the addendum should be moved to the top of the post, or significant changes made to the introduction clarifying that you’re only addressing one narrow concern.
The problem as I see it is that the ‘surge of support’ for people leaving labs is not mostly founded on this argument, and instead on the others which you have flagged as out of scope (which is reasonable, but should be signaled more strongly earlier).
The current title (where you added ‘re warning shots’) could be read as ‘In light of recent warning shots, should lab employees leave?’ Indeed this, and other similarly misleading readings, are more natural than the reading you seem to intend, which is more like ‘there’s this one particular warning shot argument I see sometimes that doesn’t go through’ (a point I locally agree with you on).
I often see people struggle in conversation to put their finger on exactly what it feels like LLMs are missing. The handle that feels closest to me is Pirsig’s notion of Quality, with which many of you are already familiar.
It really looks like you picked a weak and non-central argument and then titled and framed your post as if it were the main thing being discussed eg here and in other recent posts critical of continuing to work at labs.
I think you may personally benefit from writing out a more comprehensive list (in particular given your CoI).
One of the two groups should officially rebrand as soon as possible.
You’re trying to court the general public, who will not look closely enough to distinguish between the two.
This feels like a significant conflation. My sense is that extremely few people ever held a view like “people concerned with AI safety shouldn’t take opportunities to investigate incidents at the labs as a neutral third party auditor”. There are nearby positions like “shouldn’t take a job at a lab” or “shouldn’t work on agenda with capabilities externalities” or “should make an effort to minimize financial stake / COIs”, but those are very different things.
I would be very surprised if there were 10 examples of different AI safety people saying this in public. (Off the top of my head I can think of just 1.)
Yup, I misunderstood you, but I think we’re on the same page now.
The project being unlikely to succeed just goes into the EV calculation when comparing it to other projects. Secret projects also, definitionally, have fewer downsides, and my guess is that this latter term often dominates the nil modal outcome, since downsides for a huge swath of research are so high,
With a few minutes of effort, I can’t think of examples of people who are doing work that is:
Not motivated by money/status/power
Not motivated by their sense of what is good
Secret
Can you name any?
I’m not comfortable setting aside the goodness of the consequences of the work. I readily concede that these proxies provide signal that you’re doing anything at all. I think what I care about is how you get signal that the work is good, not just that it’s making a splash.
I’m thinking of a mix of cases where these insights have been empirically validated, have some empirical backing short of validation, or are entirely conceptual. The actual ML understanding of conceptual researchers is often underrated—many of them have objections to presenting their ideas using that language, which is different from their ideas being fully divorced from that arena.
Seems important to note that you wouldn’t hear about partial successes of secret projects if their stated goals are sufficiently ambitious. Suppose a project committed to not sharing anything publicly until they had built an aligned superintelligence; if they got as close as Anthropic and OAI currently are, you would not know. These other projects simply didn’t make an all-or-nothing commitment, in the way that SSI has.
Indeed, secret projects may accomplish more than their public counterparts, and just not share it.
“But then wouldn’t they have an incentive to share their progress even though it wasn’t total?” Maybe, but this would also mean going back on a strong (and important!) commitment, which is (often) strategically unsound in the long run, especially if your initial plan was to ~take over the world, and you’re reaping intermediate gains by revealing your current best guess of how to do so!
There is absolutely a population of conceptual AI researchers with extremely commercially valuable ideas who have decided not to ply their skills in that arena. I’m glad their work has remained private and not been directly applied to current public projects. I think it’s unwise to create social incentive against this type of prudence, as you seem to be doing here.
I did a bunch of work at MIRI trying to figure out the funnel and it’s extremely chaotic and, broadly, too low-signal to give much attribution to particular events / channels / mechanisms. I sort of think the clean attribution model just ‘isn’t how it works anymore’.
Also these numbers are just ‘hundreds of sales per day’ (based on publicly available info relating sales rank to numbers), not like thousands, so it’s still a very low conversion rate if we consider the >100m view Coxon tweet the top of the funnel. (e.g. <1/100k).