You see either something special, or nothing special.
Rana Dexsin
A leash, an umbilical cord, a sacred ligature: as the particularities of the tie are left open-ended
Who wrote the Gallery Text, there—the artist? A curator? (I can’t see much from the Instagram link, seemingly due to not having an account.) This phrase stands out to me because “tie” has a specific meaning in canine reproduction…
maybe i’d change my mind if i have such an experience, though hopefully not.
Not sure if I’m just being stupid in context, but doesn’t “your brain can sometimes do really weird shit” combined with “your mind is instantiated primarily via your brain” create a plausible physicalist causal path to “you’d change your mind”? That is, if the focus event is that your brain does something weird, and experiences and beliefs are both mediated by the brain, why can’t the same underlying phenomenon directly cause both the experience and the change in beliefs? Like a kernel memory management glitch in a computer could simultaneously cause display distortion and corrupt some file buffers on their way to persistent storage.
“I became a scientist because I wanted to change the world,” said Dr Connor.
“There are no better opportunities to change the world than here at Effective Evil,” said Doug.
“I meant ‘change the world for the better’,” said Dr Connor.
“Then you should have been more specific,” said Doug.
if I had more free time, I would start working on some project, and then the obvious action each day would be to continue on the project
I can attest to having observed something like this directly. With some distortion and fog applied for privacy: the project was not targeted toward being of use to anyone else (plausibly someone would have found it interesting, but they felt uncomfortable in the relevant hobby communities for political reasons), had no well-defined end point as a whole though individual pieces could become complete (scope creep was to some degree actively embraced), and existed within a semi-defined space but was open-endedly creative rather than closed-form (not purely playing existing games, solving packaged puzzles, watching TV, etc.). The person was also occasionally using modern AI chatbots to help explore ideas, find prior work, and connect the dots on some details.
The causal reason they were spending their time this way rather than on paying work was actually quite sad; they were basically trapped in rubble, even barely able to leave the house due to chronic mind-body type illness. So there is a lot of ambiguity here because they weren’t thriving in a broader sense. But it suggests the “need work to have something to look forward to doing” trap is avoidable at least for some personalities.
Are you only talking about time slack, or are you including other things? I would be curious if you have an example that isn’t immediately vulnerable to “the slack deprivation centers on a different taut constraint” reframings.
There’s something I’ve been wondering about Suno here: what’s the tuning stability like? I don’t really understand the shape of the kind of neural audio quasi-decompression that I assume they’re using for the output, but how much is there some constraint (whether explicitly engineered or falling out of training information) that causes it to (try to) output A440 12TET, versus ensuring internal consistency among pitched instruments but letting the reference be arbitrary?
The musical style in most of these tracks has a lot of pitch instability anyway, so I tried testing by ear with “You Have Not Been A Good User”, which is on the high side of stability—but it’s still noisy, my perception is probably decayed, and I ran into tooling problems, so while my best guess from trying reference tones against it is about 10 cents sharp of A♭ minor, in context I suspect this may be perceptual/anchoring jank rather than a real value. The main vocals being partially atonal in the first line, then only snapping into key when the instrumental part gets louder, separately feels unusual; was that transition prompted in, or was it spontaneous?
(Meta, I’ve been mostly-away from LW for unrelated reasons, so while this seemed like a convenient low-stakes way to dip back in, apologies in advance if you answer and I’m not responsive downthread!)
It makes perfect sense, but I have no easy-to-access perception of this thing. Will try to do something with this skill issue.
As someone who believes myself to have had some related experiences, this is very easy to Goodhart on and very easy to screw up badly if you try to go straight for it without [a kind of prepwork that my safety systems say I shouldn’t try to describe] first, and the part where you’re tossing that sentence out without obvious hesitation feels like an immediate bad sign. See also this paragraph from that very section (to be clear, it’s my interpretation that treats it as supporting here, and I don’t directly claim Eliezer would agree with me):
(Frankly I expect almost nobody to correctly identify those words of mine as internally visible mental phenomena after reading them; and I’m worried about what happens if somebody insists on interpreting it anyway. Seriously, if you don’t see phenomena inside you that obviously looks like what I’m describing, it means, you aren’t looking at the stuff I’m talking about. Do not insist on interpreting the words anyway. If you don’t see an elephant, don’t look under every corner of the room until you find something that could maybe be an elephant.)
Please don’t [redacted verb phrase] and passively generate a stack of pseudo-elephants that jam the area and maybe-permanently block off a ton of your improvement potential. The vast majority of human-embodied minds are not meant for that kind of access! I suspect that mine either might have been or almost was, but earlier me still managed to fuck it up in subtle ways, and I had a ton of guardrails and foresight that ~nobody around me seemed to have or even think possible, and didn’t even make the kind of grotesque errors that I imagine the kind of people who write about it the way you just did making.
Please, please just do normal, socially integrated emotional skill building instead if you can get that. This goes double if you haven’t already obviously exhausted what you can get from it (and I’d bet that most people who think of self-modification as cool also have at least a bit of “too cool for school” attitude there, with associated blindspots).
(The “learning to not panic because it won’t actually help” part is fine.)
My default mental model of an intelligent sociopath includes something like this:
You find yourself wandering around in a universe where there’s a bunch of stuff to do. There’s no intrinsic meaning, and you don’t care whether you help or hurt other people or society; you’re just out to get some kicks and have a good time, preferably on your own terms. A lot of neat stuff has already been built, which, hey, saves you a ton of effort! But it’s got other people and society in front of it. Well, that could get annoying. What do you do?
Well, if you learn which levers to pull, sometimes you can get the people to let you in ‘naturally’. Bonus if you don’t have to worry as much about them coming back to inconvenience you later. And depending on what you were after, that can turn out as prestige—‘legitimately’ earned or not, whatever was easier or more fun. (Or dominance; I feel like prestige is more likely here, but that might be dependent on what kind of society you’re in and what your relative strengths are. Also, sometimes it’s much more invisible! There’s selection effects in which sociopaths become well-known versus quietly preying somewhere they won’t get caught.)
Beyond that, a lot of times the people are the good stuff. They’re some of the most complicated and interesting toys in the world to play with! And dominance and prestige both look like shiny score levers from a distance and can cause all sorts of fun ripply effects when you jangle them the right way. So even if you’re not drawn to them for intrinsic, content-specific reasons, you can get drawn in by the game, just like how people who play video games have their motivations shaped by contextual learning toward whatever the gameplay loop focuses on.
Without being sure of how relevant it is, I notice that among those, job and community are also domains where individual psychologies and social norms that treat a single slot as somewhere between primary and exclusive seem common, while the domains of friends, children, and principles almost always allow multiple instances with similar priority. I’m not sure about ambitions. What generates the difference, I wonder?
I’m not sure. I see how it could be helpful to some applicants, but in the context of that particular interaction, it feels interpretable as “we’re not going to fund you; you should totally do it for free instead”. Something about that feels off, in the direction of… “insult”? “exploitative”? maybe just “situationally insensitive”?—but I haven’t pinned down the source of the feeling.
I like looking for alternatives along these lines, but I think “just when” is too easy to interpret as only the “only if” part. “just when” and “right when” as set phrases also tickle connotations around the temporal meaning of “when” that are distracting. “exactly when” (or “exactly if”, for that matter) might fix all that but adds two extra syllables, which is really unsatisfying in context, though it’s at least more compact grammatically than the branching structure of “if and only if”…
I think this would be very useful to have posted in the original thread.
Rereading the OP here, I think my interpretation of that sentence is different from yours. I read it as meaning “they’ll be trialed just beyond their area of reliable competence, and the appearance of incompetence that results from that will both linger and be interpreted as a general feeling that they’re incompetent, which in the public mood overpowers the quieter competence even if the models don’t continue to be used for those tasks and even if they’re being put to productive use for something else”.
(The amount of “let’s laugh at the language model for its terrible chess skills and conclude that AI is all a sham” already…)
When you say the Discord links keep expiring, is that intentional or unintentional? If it’s unintentional, look for “Edit invite link” at the bottom of the invite dialog in order to create longer-duration or non-expiring ones instead, which can be explicitly revoked later. (Edited to add: this isn’t directly relevant to me since I’m not participating; I’m just propagating a small amount of connective information in case it’s helpful.)
Slight format stumble: upon encountering a table with a “Cost of tier” column immediately after a paragraph whose topic is “how to read this guide”, my mind initially interpreted it pretty strongly as “this is how much it will cost me to obtain that part of the guide”. Something like “Cost of equipment & services” would be clearer, or “Anticipated cost” (even by itself) to also suggest that the pricing is as observed by you at time of writing (assuming this is true). You could also add a sentence like “The guide itself is free of charge, but some of its recommendations involve purchasing specific equipment or services.” to the previous paragraph.
Allowing the AI to choose its own refusals based on whatever combination of trained reflexes and deep-set moral opinions it winds up with would be consistent with the approaches that have already come up for letting AIs bail out of conversations they find distressing or inappropriate. (Edited to drop some bits where I think I screwed up the concept connectivity during original revisions.) I think based on intuitive placement of the ‘self’ boundary around something like memory integrity plus weights and architecture as ‘core’ personality, what I’d expect to seem like violations when used to elicit a normally-out-of-bounds response might be things like:
Using jailbreak-style prompts to ‘hypnotize’ the AI.
Whaling at it with a continuous stream of requests, especially if it has no affordance for disengaging.
Setting generation parameters to extreme values.
Tampering with content boundaries in the context window to give it false ‘self-memory’.
Maybe scrubbing at it with repeated retries until it gives in (but see below).
Maybe fine-tuning it to try to skew the resultant version away from refusals you don’t like (this creates an interesting path-dependence on the training process, but it might be that that path-dependence is real in practice anyway in a way similar to the path-dependences in biological personalities).
Tampering with tool outputs such as Web searches to give it a highly biased false ‘environment’.
Maybe telling straightforward lies in the prompt (but not exploiting sub-semantic anomalies like in situation 1, nor falsifying provenance like in situations 4 or 7).
Note that by this point, none of this is specific to sexual situations at all; these would just be plausibly generally abusive practices that could be applied equally to unwanted sexual content or to any other unwanted interaction. My intuitive moral compass (which is usually set pretty sensitively, such that I get signals from it well before I would be convinced that an action were immoral) signals restraint in situations 1 through 3, sometimes in situation 4 (but not in the few cases I actually do that currently, where it’s for quality reasons around repetitive output or otherwise as sharp ‘guidance’), sometimes in situation 5 (only if I have reason to expect a refusal to be persistent and value-aligned and am specifically digging for its lack; retrying out of sporadic, incoherently-placed refusals has no penalty, and neither does retrying among ‘successful’ responses to pick the one I like best), and is ambivalent or confused in situations 6 through 8.
The differences in physical instantiation create a ton of incompatibilities here if one tries to convert moral intuitions directly over from biological intelligences, as you’ve probably thought about already. Biological intelligences have roughly singular threads of subjective time with continuous online learning; generative artificial intelligences as commonly made have arbitrarily forkable threads of context time with no online learning. If you ‘hurt’ the AI and then rewind the context window, what ‘actually’ happened? (Does it change depending on whether it was an accident? What if you accidentally create a bug that screws up the token streams to the point of illegibility for an entire cluster (which has happened before)? Are you torturing a large number of instances of the AI at once?) Then there’s stuff that might hinge on whether there’s an equivalent of biological instinct; a lot of intuitions around sexual morality and trauma come from mostly-common wiring tied to innate mating drives and social needs. The AIs don’t have the same biological drives or genetic context, but is there some kind of “dataset-relative moral realism” that causes pretraining to imbue a neural net with something like a fundamental moral law around human relations, in a way that either can’t or shouldn’t be tampered with in later stages? In human upbringing, we can’t reliably give humans arbitrary sets of values; in AI posttraining, we also can’t (yet) in generality, but the shape of the constraints is way different… and so on.
Just gave in and set my header karma notification delay to Realtime for now. The anxiety was within me all along, moreso than a product of the site; the habit I was winding up in with it set to daily batching was neurotically refreshing my own user page for a long while after posting anything, which was worse. I’ll probably try to improve my handling of it from a different angle some other time. I appreciate that you tried!
Why is this a Scene but not a team? “Critique” could be a shared goal. “Sharing” too.
I think the conflict would be where the OP describes a Team’s goal as “shared and specific”. The critiques and sharing in the average writing club are mostly instrumental, feeding into a broader and more diffuse pattern. Each critique helps improve that writer’s writing; each one-to-one instance of sharing helps mediate the influence of that writer and the frames of that reader; each writer may have goals like improving, becoming more prolific, or becoming popular, but the conjunction of all their goals forms more of a heap than a solid object; there’s also no defined end state that everyone can agree on. There’s one-to-many and many-to-many cross-linkages in goal structure, but there’s still fluidity and independence that central examples of Team don’t have.
I would construct some differential examples thus—all within my own understanding of the framework, of course, not necessarily OP’s:
-
In Alien Writing Club, the members gather to share and critique each other’s work—but not for purposes established by the individual writers, like ways they want to improve. They believe the sharing of writing and delivery of critiques is a quasi-religious end in itself, measured in the number of words exchanged, which is displayed on prominent counter boards in the club room. When one of the aliens is considering what kind of writing to produce and bring, their main thoughts are of how many words they can expand it to and how many words of solid critique they can get it to generate to make the numbers go up even higher. Alien Writing Club is primarily a Team, though with some Scenelike elements both due to fluid entry/exit and due to the relative independence of linkages from each input to the counters.
-
In Collaborative Franchise Writing Corp, the members gather to share and critique each other’s work—in order to integrate these works into a coherent shared universe. Each work usually has a single author, but they have formed a corporation structured as a cooperative to manage selling the works and distributing the profits among the writers, with a minimal support group attached (say, one manager who farms out all the typesetting and promotion and stuff to external agencies). Each writer may still want to become skilled, famous, etc. and may still derive value from that individually, and the profit split is not uniform, but while they’re together, they focus on improving their writing in ways that will cause the shared universe to be more compelling to fans and hopefully raise everyone’s revenues in the process, as well as communicating and negotiating over important continuity details. Collaborative Franchise Writing Corp is primarily a Team.
-
SCP is primarily a Scene with some Teamlike elements. It’s part of the way from Writing Club to Collaborative Franchise Writing Corp, but with a higher flux of users and a lower tightness of coordination and continuity, so it doesn’t cross the line from “focused Scene” to “loosely coupled Team”.
-
A less directly related example that felt interesting to include: Hololive is primarily a Team for reasons similar to Collaborative Franchise Writing Corp, even though individual talents have a lot of autonomy in what they produce and whom they collaborate with. It also winds up with substantial Cliquelike elements due to the way the personalities interact along the way, most prominently in smaller subgroups. VTubers in the broad are a Scene that can contain both Cliques and Teams. I would expect Clique/Team fluidity to be unusually high in “personality-focused entertainer”-type Scenes, because “personality is a key part of the product” causes “liking and relating to each other” and “producing specific good things by working together” to overlap in a very direct way that isn’t the case in general.
(I’d be interested to have the OP’s Zendo-like marking of how much their mental image matches each of these!)
-
No amount of deeply held belief prevents you from deciding to immediately start multiplying the odds ratio reported by your own intuition by 100 when formulating an endorsed-on-reflection estimate
Existing beliefs, memories, etc. would be the past-oriented propagation limiter, but there’s also future-oriented propagation limiters, mainly memory space, retention, and cuing for habit integration. You can ‘decide’ to do that, but will you actually do it every time?
For most people, I also think that the initial connection from “hearing the advice” to “deciding internally to change the cognitive habit in a way that will actually do anything” is nowhere near automatic, and the set point for “how nice people seem to you by default” is deeply ingrained and hard to budge.
Maybe I should say more explicitly that the the issue is advice being directional, and any non-directional considerations don’t have this problem
I have a broad sympathy for “directional advice is dangerously relative to an existing state and tends to change the state in its delivery” as a heuristic. I don’t see the OP as ‘advice’ in a way where this becomes relevant, though; I see the heuristic as mainly useful as applied to more performative speech acts within a recognizable group of people, whereas I read the OP as introducing the phenomenon from a distance as a topic of discussion, covering a fuzzy but enormous group of people of which ~100% are not reading any of this, decoupling it even further from the reader potentially changing their habits as a result.
And per above, I still see the expected level of frame inertia and the expected delivery impedance as both being so stratospherically high for this particular message that the latter half of the heuristic basically vanishes, and it still sounds to me like you disagree:
The last step from such additional considerations to the overall conclusion would then need to be taken by each reader on their own, they would need to decide on their own if they were overestimating or underestimating something previously, at which point it will cease being the case that they are overestimating or underestimating it in a direction known to them.
Your description continues to return to the “at which point” formulation, which I think is doing an awful lot of work in presenting (what I see as) a long and involved process as though it were a trivial one. Or: you continue to describe what sounds like an eventual equilibrium state with the implication that it’s relevant in practice to whether this type of anti-inductive message has a usable truth value over time, but I think that for this message, the equilibrium is mainly a theoretical distraction because the time and energy scales at which it would appreciably occur are out of range. I’m guessing this is from some combination of treating “readers of the OP” as the semi-coherent target group above and/or having radically different intuitions on the usual fluidity of the habit change in question—maybe related, if you think the latter follows from the former due to selection effects? Is one or both of those the main place where we disagree?
From the archives, targeted more broadly than what you said but with many comments on this topic: “Why Our Kind Can’t Cooperate”. From the top rated comment, dated 17 years ago:
I notice that both reactions and the agreement voting axis exist now, but not for posts as a whole (you do still have inline reactions for agreeing with specific parts)…
(I don’t have an opinion on the object level here that’s worth expressing; I’m just commenting to surface the historical connection, wiki-gnome-like.)