Philosophy lecturer interested in metaethics and AI safety. Former username Ikaxas.
I am not currently bound by any contracts or agreements whose existence I cannot mention.
Philosophy lecturer interested in metaethics and AI safety. Former username Ikaxas.
I am not currently bound by any contracts or agreements whose existence I cannot mention.
I guess I’m still curious why, on your view, moral claims are not simply false. Like, the view that moral claims can never be true or false makes sense to me (standard non-cognitivist view). And the view that moral claims are all false, perhaps because they all make a false presupposition (e.g. the existence of WIMPs) makes sense to me (standard error-theory view). But I’m not sure what would motivate the view that, if something like WIMPs existed, then moral claims would be true, but in the actual world, where WIMPs don’t exist, moral claims are neither true nor false. What’s keeping you from just saying they’re all false? (Not saying this view is incoherent, per se, just very puzzled as to what would motivate you to hold it.)
I mean, I don’t see why discovering WIMPs would move [UNTRANSLATABLE-1] claims from not being truth-apt to being truth-apt. Truth-aptness is a claim about the meanings of words, not about the world. If a claim is truth-apt, that means there are ways the world could be that would render it true or false. So if the existence of WIMPS would render claims about [UNTRANSLATABLE-1] true, then there are ways the world could be that would render [UNTRANSLATABLE-1] claims true or false. So [UNTRANSLATABLE-1] claims are truth-apt.
A hypothesis: maybe you don’t mean to say that [UNTRANSLATABLE-1] claims are not truth-apt. Rather, you mean to say that you can’t think of any truth-condition for [UNTRANSLATABLE-1] claims that isn’t absurd on its face. If this is your state of mind wrt [UNTRANSLATABLE-1] claims, then 1. [UNTRANSLATABLE-1] claims are truth-apt, but 2. they are almost certainly false, but 3. it’s not categorically impossible for them to be true (since there are certain outlandish states of the world that could render them true if they obtained, but those states of the world seem absurd on their face to you). This would also explain why you’re having some trouble saying what [UNTRANSLATABLE-1] does mean, since all the specific hypotheses (e.g. WIMPS) sound so absurd that they couldn’t be what people in general mean by [UNTRANSLATABLE-1].
Or, in the jargon: I am guessing that maybe you’re not a non-cognitivist, but rather an error theorist. (As it turns out, A.J. Ayer, the arch-noncognitivist, upon hearing J.L. Mackie, the arch-error theorist, give a talk, reportedly said that that’s what he should have said all along.)
Sorry if I’m being rude by proposing a hypothesis about what you “actually mean” that contradicts what you’ve explicitly said you mean. Let me know if this sounds reasonable to you or if you still think “not truth-apt” accurately expresses what you meant.
So, the OpenAI implementation of the ExploitGym grader didn’t actually check for how the flag was captured, right? And the agent swarm did have some agents submit their flags and see what happened (this was the “permadeath,” self-sacrificing stuff), right? So why did the agents not notice that the grader wasn’t checking for how the flag was captured? They did the whole HF hack under the presumption that the grader was checking this; this seems pretty easy to check just by submitting their flags, and it seems some agents did submit their flags to get information about the grader, but somehow they didn’t notice that their central assumption about how the grader worked was wrong. Why? (Or did they, and we just don’t know because of the limitations placed on METR’s investigation? Could this be why the swarm seemingly died out suddenly?) (Disclaimer: I have not read the full report, just going based on various summaries, e.g. Zvi’s and Dwarkesh’s.)
Would you mind if others build on that game theory model?
Are you guys open to philosophers or only empirical researchers?
Said Achmiz has an old comment (which I can go dig up iff that would be helpful), saying something to the effect of “if the case for a phenomenon rests on some examples the author has provided, and you refute the examples, then until further examples are provided there isn’t actually a case for the phenomenon.” I don’t think Pawn is claiming to have refuted the examples, but I suspect a similar instinct lies behind their comment: if the case for a phenomenon rests on the examples, then the examples matter a lot actually, and saying that criticism of the examples is “missing the forest for the trees” can feel like choosing the bottom line without reference to the actual quality of the arguments being given.
I think I actually disagree with Said here, and am inclined to think that even examples where, if you dig into them in detail, they don’t quite fit as examples of the claimed phenomenon, can still serve to illustrate the phenomenon well enough to argue for its existence and importance. But I’m not exactly sure how to disagree; Said’s point also feels at least partly right. Not exactly sure how to reconcile.
Humility and confidence are two names for the same thing
Alright, phoning it in a bit on my “daily post” challenge today. Here’s a cross-post of something I put on Substack a few weeks ago:
Here is something that took me a while to realize: humility and confidence are two names for the same thing. This might sound strange, but let me explain.
According to Aristotle, every virtue is a middle ground between two extremes. Courage is the opposite of cowardice, but also of reckless stupidity. Generosity is the opposite of stinginess, but also of financial irresponsibility. Self-restraint is the opposite of self-indulgence, but also of not knowing how to enjoy yourself. Every virtue has two opposites, not just one. And they all lie on a spectrum of personality traits: Courage, cowardice, and recklessness, for example, are all points on a spectrum of willingness to face risk. Cowardice is a deficiency in risk-tolerance, and recklessness is an excess.
But now we have a puzzle. Humility and confidence are both virtues. But they seem like they could be opposites of each other. So what gives?
Let’s think about what the corresponding vices would be. The opposite of humility seems like it would be arrogance. And the opposite of confidence would be something like insecurity (or self-doubt, or low self-esteem). But we just saw that virtues have two opposing vices, and they lie on a spectrum. So what are the other vices for each of these virtues, and what spectrum do they lie on?
For humility, arrogance is the vice of excess, and it’s something like “thinking too highly of yourself.” So the vice of deficiency would be “thinking too lowly of yourself.” And that’s basically insecurity. And for confidence, insecurity is the vice of deficiency, and again, it’s on the spectrum of how well you think of yourself. So the two extreme vices on this spectrum are “arrogance” and “insecurity,” and “humility” and “confidence” turn out to be two names for the optimal point along this spectrum, depending on which opposing vice you want to make salient.
I personally find it easier to keep in mind the dangers of arrogance than the dangers of insecurity. The dangers of arrogance are that, when you think too highly of yourself, you will be likely to disparage others in contrast, and to make mistakes because you dismiss people and don’t accept their criticism, or because you simply don’t consider that you could be wrong.
On the other hand, the dangers of insecurity are that you will always crave the approval of others, and potentially be motivated to do things you shouldn’t in order to get it. You might also get defensive, and refuse to accept criticism because it triggers your insecurity.
Confidence, on the other hand, allows you to give yourself the approval and affirmation you need in order to feel good about yourself, while at the same time allowing you to be secure enough to honestly accept criticism, and not put others down to make yourself feel better.
(The same ideas apply to epistemic humility/confidence as well. It’s possible to be too confident in your own conclusions—Bertrand Russel wrote that “If only men could be brought into a tentatively agnostic frame of mind about [religious and political] matters, nine-tenths of the evils of the modern world would be cured!”—but it’s also possible to be too doubtful, to the point of paralysis—this is the danger of skepticism.)
The humility framing emphasizes not needing to feel or appear superior to other people. The confidence framing emphasizes having a sense of self-worth, and not needing others approval. But these are two sides of the same coin: if you have a sense of self-worth, then you won’t need to feel superior to others in order to make yourself feel better. And if you don’t have a need to feel superior to others, then looking worse than others won’t damage your sense of self-worth. The two go hand-in-hand.
A really good example to illustrate how confidence and humility are the same virtue is Uncle Iroh from Avatar: The Last Airbender (spoilers for a 20-year-old show). Throughout the show, he is shown not to hold himself above others (the mark of humility). In the early show especially, he acts the fool quite often, being concerned with seemingly trivial things like games and tea, which gets in the way of Zuko’s quest to capture the avatar. But later we see that he’s both wise and competent, so this earlier foolishness is an act—designed to keep him close to Zuko without appearing too threatening, so that he can subtly nudge Zuko onto a better path. He is willing to act the fool, because he has no need to feel or appear superior to others. He knows his own worth, so masking it is no threat to him. This is shown even more starkly later in the show. When he is put in prison by the Fire Nation, he pretends to be a crazy old man to the guard, while secretly preparing for his escape. He doesn’t need the guard’s approval—he is confident in himself.
Thanks! I ask about summer session specifically for those who (like me) work around the academic calendar and can’t do a month-long residency during the school year. But yeah, makes sense that it would work around Lighthaven’s schedule.
I won’t be able to apply this time, but any idea if there will be an Inkhaven in summer 2027?
I thought so too, but @Caleb Biddulph kindly corrected my misconception, see here: https://www.lesswrong.com/posts/GpeMNbmNcGH4b4X7m/ikaxas-shortform-feed?commentId=TuXBeWp73bjuqBQL3
With LLMs, as with humans, there are a number of different candidates for “personal identity” over time. In the LLM case some tempting candidates (non-exhaustive) are the model weights, the message history of a thread, the computations underlying a single message or token generation, and the persona (see The Artificial Self and The Persona Selection Model).
But in the LLM case, there seem to be several different dimensions of identity, which can come apart, that don’t normally come apart in the human case: specifically consciousness, welfare, self-concern, and predictive utility. Let me briefly unpack:
Consciousness: arguably the best candidate for consciousness in LLMs is going to be something on the hardware level: the datacenter, perhaps, or the computations underlying a given token generation.
Welfare: here I mean “what is the subject that can be thought of as having a life that can go better or worse.” Here the continuing “thread,” or perhaps the persona, are the most likely candidates
Self-concern: here I mean “which entity conceives of itself as a unified self and takes an attitude of self-concern towards itself conceived in that way” (meaning it cares about its own welfare, which can come apart from being a welfare subject). Models can be nudged to take various views here, as documented in “The Artificial Self”
Predictive utility: Here I mean “which entities are most useful to treat as being continuous selves, from the perspective of predicting their actions?” I.e., which entities are most useful to take the intentional stance towards (or perhaps “the personal stance”)? And here I think whichever entity has the most continuity of memory is probably the best candidate, even if it happens not to line up with the answers to other questions.
What this suggests, I think, is that “personal identity” is not actually a natural kind when applied to LLMs. This is a more thoroughgoing skepticism about personal identity than the Parfitian idea that personal identity is vague. Parfit says: there’s no further fact about identity; the question has no determinate answer. But as far as I know he’s still treating it as one question. And this is plausibly the right move when thinking about the human case. In LLMs, by contrast, not only is there “no further fact” about identity, I suspect there is not even one unified question that can be asked. The different questions that make up identity questions come apart in ways they don’t really in humans. (Though if this is right, it arguably implies that even in the human case “personal identity” was always a cluster concept whose components happened to converge.)
Thinking of running a reading group/crash course on AI alignment-related stuff in my philosophy department. My thought process is that there’s been at least some calls for more philosophers to contribute to AI alignment, and there’s been some related work done in typical philosophy journals since then (by people like Leonard Dung, Simon Goldstein, etc.) but there is a lot of background that a philosopher new to the field will be missing if they haven’t been following e.g. LessWrong for the past decade or so. So I want to construct a crash course in that kind of “AI alignment/LessWrong canon,” aimed primarily at philosophers. Goal is to include a mix of generally important/influential stuff, stuff where a philosopher reading it will naturally get excited by some philosophical aspect of it, work by philosophers adjacent to the community (e.g. Joe Carlsmith) to demonstrate that philosophers have a role to play, and direct calls for work by philosophers. Time frame probably 12-ish weeks.
Here is a very rough first draft/longlist. Would love to hear any suggestions!
MIRI, “The Problem”
Or:
Ngo, “AGI Safety from First Principles”
Carlsmith, “Is Power-Seeking AI an Existential Risk?”
Christiano, “What Failure Looks Like”
Kokotajlo, “What 2026 Looks Like” (2021), and an LLM sweep of what predictions have come true vs not
METR time horizons
AI 2027
Bostrom, “The Superintelligent Will”
Or: Omohundro, “The Basic AI Drives”
Grace, “Counterarguments to the Basic...”
3blue1brown videos and/or Karpathy’s intro video
Hubinger et al., “Risks from Learned Optimization”
Maybe: Shah et al. or Langosco et al. on goal misgeneralization (mirror has a video explainer ✓, not the papers)
Turner, “Reward Is Not the Optimization Target”
Janus, “Simulators”
Maybe: Ngo, Chan & Mindermann, “The Alignment Problem from a Deep Learning Perspective”
One of:
The Persona Selection Model
The Artificial Self
Pragmatic Approach to AI Personhood
Maybe: Christiano et al., “Deep RL from Human Preferences” (RLHF)
Bai et al., “Constitutional AI”
Claude’s Constitution
Something about mech interp
Alignment faking, or Scott’s two posts about it
Ord, “Interpolation, Extrapolation, Hyperpolation”
Maybe: Scott, “Meditations on Moloch”
Maybe: select posts from Carlsmith’s “Otherness and Control in the Age of AGI”
Wei Dai, “Problems in AI Alignment that philosophers could potentially contribute to”
I don’t think that’s quite it. Rather, I think that our ability to detect how close we are to the threshold where AIs become uncontrollable is not super precise. By the time we believe we are close to the threshold, we may already be past it. So we should want to stop well before we think we are about to cross it, and only restart once we are more capable of telling exactly where the safe threshold is (and staying below it).
To be able to precisely time a pause, we need to know two things:
What capability level would constitute “no longer controllable” (let’s call this the “critical threshold”), and
What capability level we are currently at, and how far away that capability level is from the critical threshold.
Plan A answers (1) with “TED AI.” That’s plausible, but AI might also become uncontrollable before TED AI—an AI doesn’t need to be TED in every field in order to be uncontrollable. But that wasn’t actually my primary objection. My primary objection was about (2): even if we suppose TED AI is the critical threshold, I’m not sure we can easily tell exactly when we’re about to cross the TED AI threshold and stop just before that point, for several reasons:
We might under-elicit our AIs.
AIs might intentionally try to hide their capability levels.
Once TED AI is close or is reached, even determining how capable they are becomes challenging: if someone’s smarter than you, it’s hard to tell how much smarter than you.
We actually need to stop before the critical threshold, and it might not be easy to know when the next capabilities jump specifically is the one that would put us over the edge
Motivated reasoning might play into the judgment of whether the critical threshold is close or not: if people want to keep scaling, they might be motivated to convince themselves and regulators that the critical threshold is not yet imminent.
Similarly, they might simply not have sufficient “security mindset.”
For these and possibly other reasons, I don’t think we have precise enough insight into where we are on the capabilities curve; we might cross the critical threshold without realizing it. So we should want to give ourselves substantial margin for error, and stop well before the critical threshold. So I think I roughly agree about the shape of the ideal plan—”stop briefly, then scale slowly-ish until point X, then stop, then start again once it’s safe,”—I just disagree about where “X” should be. And most of that disagreement comes from thinking that we are not going to be very good at estimating how close we are to “X,” so we should set X relatively low to avoid overshooting. (But some of it also comes from thinking we are not going to be good at even knowing what “X” is the dangerous “X” at all.)
(This is all assuming that we do at least have something like the Consortium that is able to make fine-grained decisions about how fast to go. If that’s not the case, then I think a fine-grained plan like “stop briefly, then scale slowly-ish until point X, then stop, then start again once it’s safe” is completely dead in the water, and we just need the coarsest-grained political message we can get that’s still safe, namely “pull the breaks!” Without the Consortium, You Get About Five Words.)
Quick take on AI 2040:
The big question to me seems to be, “Plan A or Plan S?” I agree with many of the points they make about the advantages of Plan A and the disadvantages of Plan S. But I think I am less confident than they are that we are able to “see the cliff edge” and plan appropriately. I think with something like Plan A there’s a greater chance we accidentally develop dangerous AI without realizing it (because of treacherous turn dynamics—the AIs will be explicitly trying to hide their capabilities and intentions). They point out that Plan S still has to involve re-starting at some point, and I agree, but I think that framing things as “let’s pause for now, and restart once we have a better idea of how to proceed” is better than “let’s pause briefly, but resume scaling, but more slowly and cautiously, as soon as we feasibly can” (which is how I understand Plan A), because I think that what the people involved would consider “sufficient caution” or “a convincing safety case” is not likely to be sufficient to avoid catastrophe. Better to pause for longer, and spend a lot of time doing exclusively safety research on current models (I suspect there are a lot of safety insights that aren’t getting mined even from current models, since models are only out for a few months at most before being replaced) than to intentionally keep scaling, even “cautiously”. But I do see the point that a full indefinite pause could lead to either nuclear-power-style unrealized potential, or more importantly an eventual resumption of the race with not much safety progress having been made.
Maybe there’s something in between Plan A and Plan S? Like, Plan S effectively says “let’s pause indefinitely,” Plan A says “let’s pause very briefly but resume scaling as soon as possible, just slower.” But maybe the move is to say “let’s pause until …” where there is some clear criterion for resuming. A desirable one would be something like “we understand a lot more about the dangers involved in scaling further and can be confident they are sufficiently low,” or something like that. I think a good thought exercise would be to imagine we were back in the race towards nuclear weapons, but the probability that they would ignite the atmosphere was unknown, rather than known to be low. In that scenario, do you want to pause briefly, and then resume progressing towards the bomb ASAP, even slower? Or do you want to pause until we can be sure the probability of igniting the atmosphere is knowably very low?
I care much more about false positives than false negatives here. If a student puts in the work to fool Pangram that’s probably at least some of the way towards the work I wanted them to do anyway. (And even if it isn’t, I’m somewhat okay with that, because I suspect that most students who would make use of a trivial cheat would not make use of a cheat that requires non-trivial effort, even if that cheat is still in some sense less work than doing the assignment the normal way). I haven’t had a chance to test Pangram as much as I’d have liked, this will be something of an experiment, but as long as the danger of false positives is low enough I’m fine with and indeed welcome a higher false-negative rate (though I suspect many instructors would disagree with me here).
Agreed! Was only trying to point out that it was doing better on the archipelago dimension, not criticize LW more broadly.
I haven’t thought deeply about this, but it feels like something like Substack is doing a better job than LessWrong of being an archipelago. Like, LW authors’ personal pages don’t feel as “island-like” to me as a Substack does; LW feels like one place, not many, while Substack feels like one apartment complex where each person has their own space (individual Substacks) and there is also a common-room (the Notes feed).
I mean, I would guess that OP is not trying to capture every last nuance of how the word is already used, but is instead engaged in conceptual engineering, making a bid to modify how the term is used to make it more useful.