Rat adj since forever, actually active again since AI alignment is a real world problem now. I am not a programmer or a mathematician but am nevertheless drafted to teach serious professionals in serious roles what AI is and how/if/when they should use it bc they couldn’t find anyone else i guess.
Loki zen
This isn’t especially surprising to me, and some of it’s true! Whether or not agentic/rogue AI is a danger to most people, the fact remains that perfectly well-aligned* AI following current trends and aspirations in implementation has basically all the risks that they are talking about, and the fact that “AI risk” types rarely talk about those things does make them seem out of touch /like they’re not taking this seriously to a lot of people.
*in the way that ‘aligned’ is typically, in practice, used—meaning that it will do what the people paying for it and operating it want it to do.
yeah I mean that’s the point—we’re not where the money is
I also appreciate that they removed the reference to Anthropic’s revenue as a goal for Claude from the previous leaked version which included “Claude acting as a helpful assistant is critical for Anthropic generating the revenue it needs to pursue its mission.”
text to more or less this effect is still in there. “Claude is also central to Anthropic’s commercial success, which, in turn, is central to our mission.” I don’t think that’s meaningfully different.
absolutely people are less likely to want a model that can whistleblow against your wishes—regardless of whether or not it is a screwup! but that’s a commercial consideration dressed up as “alignment”, which is one part of this that I really hate.
healthcare)
Aside from purpose-built image models in radiology this is IMO, in the short term at least, absolutely one of the riskiest uses of AI. It’s a situation where lives, in the moment, depend on getting and recording absolutely accurate information, and people are right now schmoozing the suits at your local hospital trying to get them to make clinicians use AI tools that aren’t fit for purpose to “summarise” evidence and make patient records. This has the potential for harm in the direct sense and also the potential to severely harm peoples’ trust in medical institutions, which is a very dangerous thing to have happen.
the above seems a potential likely cause of the see-sawing you’re talking about too
AIs will be evaluated, inspected, and selected by us, and their behavior will be determined directly by our engineering.
I think the “us” here is meaningfully and importantly different from the “us” in the rest of this paragraph. This “us” may include the majority of people reading this article, but it doesn’t include me or vast the majority of humans, and that has a meaningful impact on how comforting this prediction is able to be.
You appear to be attempting an explanation of why males would evolve to sexually discriminate with respect to breasts.
This is a separate question from why breasts evolved. Are you suggesting that the existence of an alternative mechanism for male interest in breasts factors somehow into how we should weight the sexual selection hypothesis?Or have you just misunderstood the sexual selection hypothesis? There is no genetic incentive to evolve a visible reproductive use-by date, however convenient it might seem from a top-down species perspective.
(I also think you’re stating without much justification a difference in the universality of the visual effects of ageing on the skin vs the breasts that simply might not exist.)
If you feel better, healthier, and/or have better biomarkers, when you decrease the amount of X in your diet, maybe you would benefit from cutting it down to zero.
would argue this is specifically not true about nutrition, because all else being equal more variety usually = better nutrition
Thanks for linking me. But no, I don’t really find that post illuminating with respect to the question that I have here? I see that you are making a case for why you find it reasonable to be concerned with what you are concerned with, but unless I’m missing something it still feels like at every turn the reasoning is from how you have chosen to define entirely hypothetical things that you’re pretty sure somebody is going to invent some day.
“Me: As it happens, the threat model I’m working on is not LLMs, but rather “brain-like” Artificial General Intelligence (AGI), which (from a safety perspective) is more-or-less a type of actor-critic model-based reinforcement learning (RL) agent. LLMs are profoundly different from what I’m working on. Saying that LLMs will be similar to RL-agent AGI because “both are AI” is like saying that LLMs will be similar to the A* search algorithm because “both are AI”, or that a frogfish will be similar to a human because “both are animals”. They can still be wildly different in every way that matters.”
How would you respond to the critique that this basically amounts to saying “I’m not interested in or saying anything about things that exist, but this thing that AI X-Risk types made up because it’s the most worrying hypothetical to talk about sure is a worrying hypothetical.”
[glibly phrased but meant in a spirit of genuine curiousity because I still don’t really understand why people care about AI risk but lack any interest in stuff that actually exists right now]
It may well be. It’s been my observation that what distracts/confuses them doesn’t necessarily line up with what confuses humans, but it might still be better than your guess if you think your guess is pretty bad
Yeah, we do.
but they’re not agents in the same way as the models in the thought experiments, even if they’re more agentic. The base-level thing they do is not “optimise for a goal”. We need to be thinking in terms of models that are shaped like the ones we actually have instead of holding on to old theories so hard we instantiate them in reality
I don’t know how you “solve inner alignment” without making it so that any sufficiently powerful organisation can have an AI of whatever level we’ve solved that for that is fully aligned with its interests, and nearly all powerful organisations are Moloch. The AI does not itself need to ruthlessly optimise for something opposed to human interests if it is fully aligned with an entity that will do that for it.
TheAIcorporation does not hate you, nor does it love you, but you are made out of atoms which it can use for something else.
my take is that they haven’t changed enough. People often still seem to be talking about agents and concepts that only make sense in the context of agents all the time—but LLMs aren’t agents, they don’t work that way. if often feels like the Agenda for the field got set 10+ years ago and now people are shaping the narrative around it regardless of how good a fit it is for the tech that actually came along.
Good post, and additional points for not phrasing everything in programmer terms when you didn’t need to.
more provocative subject headings for unwritten posts:
I don’t give a fuck about inner alignment if the creator is employed by a moustache-twirling Victorian industrialist who wants a more efficient Orphan Grinder
Outer alignment has been intractable since OpenAI sold out
1. many commercial things actually are just better (and much more expensive) than residential things. This is because they are used much more by people who are less careful with them. A chair in a cafe will see many more hours of active use over a week than a chair in most peoples’ homes!
2. a huge amount of residential property these days is outfitted by landlords—that is, people who don’t actually have to live there—on the cheap, and with as little drilling into the walls (affecting the resale value) as possible.
inasmuch as personalised advice is possible just from reading this post (and as, inter alia, a pro copyeditor), here’s mine—have a clear idea of the purpose and venue for your writing, and internalise ‘rules’ about writing as context-dependent only.
“We” to refer to humanity in general is entirely appropriate in some contexts (and making too broad generalisations about humanity is a separate issue from the pronoun use).
The ‘buts’ issue—at least in the example you shared—is at least in part a ‘this clause doesn’t need to exist’ issue. If necessary you could just add “(scripted)” before “scenes”.
Did someone advise you to do what you are doing with LLMs? I am not sure that optimising for legibility to LLM summarisers will do anything for the appeal of your writing to humans.
is it just me or is this very obvious AI generated text