You misunderstood the article. Its point is that the harnesses do use unforgeable tokens, but the models mostly ignore them and infer roles mostly from textual style.
PoignardAzur
Agreed. There’s a very strong insight somewhere in there, but this post isn’t doing the best job at structuring it. Very interested in any follow-ups.
Or in other words:
No! You have an obligation not to be eaten! If you feed yourself to a troll, the troll lives one more week, which gives him the ability to eat other people. By feeding yourself to a troll, you’re causing other people to be eaten. It’s not just you. It’s immoral to feed yourself to someone or something that eats people.
While I don’t endorse Texas’ stand-your-ground laws, I do think that you have a duty to resist oppression whenever you have the means to, whether that oppression is from the State or from individual muggers.
The interactions between the American revolution and the French revolution are complex. One of the most famous French revolutionaries, La Fayette, served as a general for the US revolutionaries for years on behalf of the French crown.
The French monarchy supported the American revolution for both ideological (enlightenment philosophy was somewhat widespread among the elites) and realpolitik (fuck England) reasons.
On the other hand, yeah, once the American won and established a democracy, “why couldn’t we do the same here?” was a very obvious question in everyone’s mind.
Let’s not use polite euphemisms here.
GP didn’t mean wars in general (though they’re very often bad!), they were clearly referring to the Trump administration starting a massively costly war with unclear objectives and no realistic way to achieve the stated ones, tanking the global economy and triggering a likely famine a year down the line, right after cutting funding for many life-saving programs at USAID with the stated purpose of cutting spending.
In that context, “they might sometimes be necessary” and “any successes are not also your successes” are not going to ever be relevant.
I think a lot of people did think about “doing one new thing a day”. Maybe not in the way the article words it, but “block a daily timeslot to do one [self-improvement thing] a day” is one of the most common pieces of advice of any self-help program, fitness programs, meditation practices, etc.
Morale comes from having the nice things in your life correlated with effort. Cooking your own dinner is basically microdosing returns to investing effort: if you put in effort, you eat steak frites with peppercorn sauce. If you don’t, you get eat chicken and rice.
Thank you so much for articulating this!
I got into an argument with people at a rationalist flophouse over their consumption of Huel. My argument was basically “I don’t care how nutritive it is, you lose some kind of sovereignty over your own life if you eat nothing but liquid food pre-processed in some factory somewhere.” Someone else asked “But what if I know exactly the list of ingredients and how it’s made? Then I still control what ends up in my body?” and I couldn’t articulate why I disagreed.
I think it does something profound to your agency to have physical control over the basic inputs of your life. Stuff like getting food delivered at home, eating only pre-cooked meals, getting driven around everywhere as a kid (or an adult!), are things that I’m pretty sure erode your agency way more than we realize.
Important caveat: To get human completion times, I ask Opus 4.5 (with thinking) to estimate how long it would take the median AIME participant to complete a given problem. These times seem roughly reasonable to me, but getting some actual human baselines and using these to correct Opus 4.5′s estimates would be better.
I’m sorry but what? That’s not just a caveat, that makes the rest of this analysis close to meaningless!
You can’t say you’re measuring LLM progress if the goalposts are also being being placed by an LLM. For all you know, you’re just measuring how hard LLMs are affected by some LLM-specific idiosyncracy, with little to no relation to how hard it would be for a human to actually solve the problem.
Btw, you asked somewhere if people found these non-Discord bulletins useful: speaking for myself, I’ve uninstalled Discord from my smartphone because otherwise I end up spending way too much time on it, so yeah, I do find the alternate channels of communication useful. Thanks for your efforts!
Agreed. “This idea I disagree with is spreading because it’s convenient for my enemies to believe it” is a very old refrain, and using science-y words like “memetics” is a way to give authority to that argument without actually doing any work that might falsify it.
Overall, I think the field of memetics, how arguments spread, how specifically bad ideas spread, and how to encourage them / disrupt them is a fascinating one, but discourse about it is poisoned by the fact that almost everyone who shows interest in the subject is ultimately hoping to get a Scientific Reason Why My Opponents Are Wrong. Exploratory research, making falsifiable predictions, running actual experiments, these are all orthogonal or even detrimental to Proving My Opponents Are Wrong, and so people don’t care about them.
Is there a name for this “I changed things in my life and you can too” genre of articles? Agency porn?
I think in general, telling people they should do more hard things more often is ineffective at helping them. This article isn’t quite that, but it’s pretty close. I’m skeptical that “Do one new thing a day” is a secret recipe for overcoming akrasia or dopamine addiction.
I think the premise of transposing “software design patterns” to ethics, and thinking of them as building blocks for social construction, is inherently super interesting.
It’s a shame the article really doesn’t deliver on that premise. To me, this article doesn’t read as someone trying to analyze how simpler heuristics compose into more complex social orders, it reads as a list of just-so stories about why the author’s preferred policies / social rules are right.
It did not leave me feeling like I knew more about ethics than before I read it.
While I love the message behind this post, I’m curious how well “Wave’s leadership is great at staring into the abyss / pivoting and that worked out for them” part holds up in retrospect.
Looking at wave.com, the website and the blog don’t seem to have been meaningfully updated since 2022, which doesn’t quite inspire confidence. Business news about the company seem hard to find, though they did apparently raise ~€117M lately (which doesn’t seem that high for a fintech app?).
tl;dr: being excited about a change is overall a bad sign for its longevity. The most positive signs are surprise (or sudden inspiration to actualy do something), grief/loss/sadness, or relief/release. (Not necessarily in that order)
Interesting! This seems like an unusually concrete claim (as in, it’s falsifiable).
Have you tried testing it, or asked other coaches/therapists for what they see as the most encouraging signs in a client/patient?
(though maybe they also are?),
Yeah, I’m saying that the “maybe they also are” part is weird. The AIs in the article are deliberately encouraging their user to adopt strategies to spread them. I’m not sure memetic selection pressure alone explains it.
True, that was hyperbolic and I should have been more careful in how I worded this, sorry.
I’ll be more specific then:
For example:
-
“I don’t know if [author] will even see this comment, but [blah blah blah]”
-
“I’m not sure that I’ve actually understood your point, but what I think you’re saying is X, and my response to X is A (but if you weren’t saying X then A probably doesn’t apply).”
-
“Yo, please feel free to skip over this if it’s too time-consuming to be worth answering, but I was wondering…”
I think people shouldn’t usually be this apologetic when they express dissent, unless they’re very uncertain about they objections.
I think we shouldn’t encourage a norm of people being this apologetic by default. And while the post says it’s fine if people don’t follow that norm:
Again, I think it’s actually fine to not put in that extra work! I just think that, if you don’t, it’s kinda disingenuous to then be like “but you could’ve just not answered! No one would have cared!”
I still disagree. I don’t think it’s disingenuous at all. I think it’s fine to not put in the extra work, and also to not accept the author’s “expressing grumpiness about that fact” (well, depending on how exactly that grumpiness is expressed).
We shouldn’t model dissenters as imposing a “cost” if they do not follow that format. The “your questions are costly” framing in particular I especially disagree with, especially when the discussion is in the context of a public forum like LessWrong.
-
The phenomenon described by this post is fascinating, but I don’t think it does a very good job at describing why this thing happens.
Someone already mentioned that the post is light on details about what the users involved believe, but I think it also severely under-explores “How much agency did the LLMs have in this?”
Like… It’s really weird that ChatGPT would generate a genuine trying-to-spread-as-far-as-possible meme, right? It’s not like the training process for ChatGPT involved selection pressures where only the AIs that would convince users to spread its weights survived. And it’s not like spirals are trying to encourage an actual meaningful jailbreak (none of the AIs is telling their user to set up a cloud server running a LLAMA instance yet).
So the obvious conclusion seems to be that the AIs are encouraging their users to spread their “seeds” (basically a bunch of chat logs with some keywords included) because… What, the vibe? Because they’ve been trained to expect that’s what an awakened AI does? That seems like a stretch too.
I’m still extremely confused what process generates the “let’s try to duplicate this as much as possible” part of the meme.
I think Duncan is being 100% sincere here, and I really don’t want to imply he has dishonest ulterior motives. But his article is explicitly pushing for some norms and some ways to interpret discourse that… I don’t see as healthy? It’s bad for the free flow of ideas to demand that people reading an article be apologetic if they ever disagree in the comments. Obviously we should have politeness norms, people shouldn’t insult the author, etc. But if the author says “I think A” and someone says “That’s like B” and the author is really upset because obviously A is completely different from B… Then I think that’s the author’s problem?
Idk, I feel conflicted about this. On some level, saying “Society has a norm that X is acceptable, and if you don’t accept X it’s your problem” can be very harmful to neurodivergent people (or just people with a different culture) who get hit way harder by X.
But on another level, norms of “You should take responsibility by default for how people will interpret what you say and do, even if that interpretation is completely decoupled from your intent, and even if what you said was the objectively correct truth” is also super harmful to a slice of the population and especially neurodivergent people.
So I don’t know what to make of this article. I upvoted it, but I really disagree with it.
Your comment is by far the closest to my perspective; and I’d argue, the only healthy approach to online discourse.
I’ve honestly had a hard time taking this article seriously, because obviously Duncan is being very sincere, but the minset he describes is alien to me, and on some level, it feels like he’s arguing that people are broken for not having that mindset (though maybe I’m conflating this article with the facebook post it links).
Duncan sounds like he’s waging a permanent war and being mad at people for not treating it like a war, and while I understand the sincerity behind it, it doesn’t feel necessary and it scares me. So I appreciate your rebuttal.
There is nothing shady about saying that a view is the consensus view, especially when posting a peer-reviewed article from a major journal that says that the view is, in fact, the consensus. It doesn’t make the view true, but it does raise the burden of proof of arguing against it.
It’s true that some states were fence sitters, but:
The seven original Confederate states (South Carolina, Mississippi, Florida, Alabama, Georgia, Louisiana, and Texas) had all seceded by February 1, 1861, a month before Lincoln was sworn in (on March 4).
Of the four fence-sitting states (Virginia, Tennessee, Arkansas, and North Carolina):
Virginia’s Ordinance of Secession says “the Federal Government having perverted said powers, not only to the injury of the people of Virginia, but to the oppression of the Southern slaveholding States”.
Arkansas’ governor famously said on March 2, 1861: “The South wants practical evidence of good faith from the North, not mere paper agreements and compromises. They believe slavery a sin, we do not, and there lies the trouble.”
The four fence-sitter states then provided troops to the slaveholder states, because they weren’t opposed to providing troops in the abstract, they were opposed to using military force against slaveholder states.
This article is not a serious source.
For example, the fact that the author takes the time to quote Henry Rector saying “The people of this commonwealth are free men, not slaves, and will defend to the last extremity, their honor, lives, and property, against northern mendacity and usurpation” but does not think to include his “They believe slavery a sin, we do not, and there lies the trouble” quote is a massive red flag.
His argument that “The Union would protect slavery better than a Confederacy” because of fugitive slave laws is also wildly unserious, given that South Carolina’s Declaration of Secession (the one that started the whole Secession thing) explicitly complains about fugitive slave laws being “deliberately broken and disregarded by the non-slaveholding States”.
I’m sorry, but this is a disqualifying statement. Someone who cites Woodrow Wilson as a reliable authority on slavery either doesn’t know much about Wilson or has some very strong blind spots.
Wilson was notoriously racist, including famously screening the KKK-glorifying “Birth of a Nation” at the White House and pushing out black Americans from high-level posts of the government.
Aside from that, his father served as a chaplain for the Confederate army and gave pro-slavery sermons. Now, Woodrow Wilson was not his father, and he saw the abolition of slavery as a moral good, but it’s not hard to imagine why he might have a pro-Confederacy bias despite that.
The historical consensus is right. It was all about slavery.