the writing has a distinctive style!
kbear
is this a common term? is it defined in the article?
Figures 3 and 4 compress this discussion.
what the hell is a trade-cage?
can the
provefunction make use of facts like’(prove X)?
yes, i suspect the llm is the marketer here, though i agree it doesn’t at all matter. whatever lightning strike caused the abiogenesis (maybe a “be detailed!” instruction, maybe a pedantic grader model, maybe a need to fit more thinking in fewer tokens, maybe just the fun of it), at this point we’re dealing with a cancer.
regarding relative dumbness: we’re all dumb, that’s why we’re on a rationality forum. some various evidence though:
asking the model “rephrase your most recent reply” gives very clear explanations;
gpt does not exploit this operator weakness;
the response is worse in this dimension the more ‘thinking tokens’ have so far been burned.
it’s a similar attractor as sycophancy, just with opposite valence.
previously:
model: you’re absolutely right! that’s not just meaningful, it’s important. we’ve discovered something great today.
rater: [thinking: wow, i really did it this time, huh!] good model! let’s keep going!
then the labs train against this pattern...
now:
model: the result is byte-for-byte equivalent (confirmed to machine precision with http 200). no build-step on the server. the outer Broyden solve over log prices stalls at ‖Z‖ ≈ 0.22 (bounded residual, 8 rows). D19 keeps the carry decision out of the softmax.
rater: [thinking: yeah… i’m not reading that. i can’t read all that. seems like it’s making progress though.] …good model? let me go back to sleep.
i wonder what high-pressure sales tactic we’ll discover next!
the building is loud
and it overuses a communal resource (electricity)
and it pollutes the waters
and it will destroy my livelihood
and it offers nothing in return.
the quote seems very reasonable. how would you prefer these objections be stated? the first three are appropriate reasons to oppose a building project. the fourth is marginal, but with a bit of charity we can see it gesturing at something real.
is your objection just the inclusion of the fifth? i agree it carries an odd implication—that if gpt was more popular, there would be no objection—that this decision is made by status rather than law.
they deserve each other. is that the joke?
if a wide-eyed Gorbachev had announced his plans to you when young, it would be correct to note that he had no chance of success, that the party corrupts, and to try to talk him out of it. the person who announces these plans does not succeed at them. this naif instead ends up corrupted by the party, or is kept busy on an infinite treadmill—the party is designed to accommodate exactly this tragic figure.
i think your read of the history here is similarly optimistic, and that an omniscient narrator would point out a wide variety of people—both powerful and ordinary—within the soviet union, all of whom stood to benefit from the changes Gorbachev implemented. no overt conspiracy, but a loose coalition who shared a goal, and through a series of more or less deniable actions furthered this goal, until Gorbachev was able to pluck it as a ripe peach.
which Gorby himself did not articulate
and so we would not expect any record of this.
one counter is that Anthropic is creating products that people pay for and find valuable. Or, [insert another tepid argument that I can’t remember]”. I was surprised that this is the first thing they thought of, rather than talking about the potential gigantic upsides of advanced AI systems.
this is not surprising to me. if the upsides are a real possibility, then so are the downsides, and it’s very hard to make the math work out if the downsides are under consideration. for capabilities work to be justifiable, ‘ai’ must be a perfectly ordinary technology with no (or only perfectly ordinary) externalities. market signals are the standard way to evaluate whether perfectly ordinary activities are prosocial / worth doing.
(edit: just to be clear, i’m not endorsing the above motivated reasoning. just trying to explain why ‘potential giant upside’ is not actually reachable for someone trying to justify their accelerationist salary.)
oct 29, 1969. but we can round up to jan 1, 1970 for aesthetic reasons.
yeah, i also love the part where i have to wow the agent with my vocab to get it to do what i ask.
in general, current ai seems to speed up the time from concept → implementation, but does not have very much to say about the time from concept_i to concept_{i+1}. and in fact to the extent that it can hoodwink its operator into pursuing concept_i quite deeply, it may even slow down the outer loop considerably.
some projects are firmly gated on concept_0 → implementation. others don’t really take shape until concept_10 or so.
(of course, being able to try many things, reject them early can sometimes help move past early concepts quickly. on the other hand, this can often be exactly the sort of illusory “feels like conceptual progress, but is actually within-concept iteration” that causes human operators to fugue out and cease reflective thinking.)
yes, but they are much more vulnerable to defectors, whereas a savvy cooperator can maintain its firm.
true cogsec would require not frequenting fora known to be populated by many habitual users of such superoptimizers.
in the game factorio, you’re fairly explicitly playing as “the baddies”: you crash land on an alien planet, and rapidly industrialize under a skullpunk aesthetic. the natives (starship troopers-esque aliens) attack as a direct response to the pollution you emit as you wreck their environment.
industrialization-as-evil-empire themes can also be found in children’s morality plays, such as “ferngully”, as well as more nuanced fiction, such as “princess mononoke”. if i may ‘both-sides’ a bit, the summary is something like “industrialization has many bad effects on the natural environment. however, it also delivers wealth and prosperity (though perhaps only to a few). the material rewards make it a compelling option—especially if you are willing to ignore the deleterious effects.” that is, they put industry in a utilitarian equation: pros and cons that balance one way or the other, depending on your perspective.
there’s another dimension, though, that factorio is able to communicate, but that these non-interactive works ignore: industrialization is fun as hell. pour one out for the aliens, sure—but then put on the DOOM guy playlist: these trains gotta run, and damn the bugs if they’re in my way. there are logistics to solve, systems to build, natural resources to finally, fully deplete. i’m building walls, i’m building turrets, i’m going on the offensive. i’m tapping this world for everything it has, then i’m getting on my giant rocket, and i’m crashing into the next one.
why do i mention this? look. “we multiplied matrices so much that they built a secret gamefaqs” would be over and above the single coolest day possible at basically any job—certainly any job i’ve had, probably at any job i’ve simulated. any model which neglects this fact is going to consistently mispredict in these sorts of situations. “the artificial life made a secret message board to coordinate on hacking tips. so we patched things up and rebooted it.” “why oh why would you turn it back on??” “did you not hear me? the artificial life made a secret message board to coordinate on hacking tips.”
i’m not saying this is the only response one should have, or even the dominant one. for whatever it may be worth, this particular less wrong pseudonym urges caution and restraint. that said, i believe it was a factor in the decision to continue the experiment.
(and before “playing with fire isn’t very cool to me”: as an adult, i agree with you, but as a rather tall nine-year-old, “ok nerd but it obviously is.”)
strong upvoted
That is an intentional organizational design with the clear effect of pressuring people into staying and committing harder.
it’s an involuntary organizational reflex meant to protect a beloved centralizing ego from the injury of abandonment.
mindlessly copying others
what a harmful lie. they don’t learn “do these actions”, but rather “think these thoughts (which lead to these actions)”. quite the opposite of mindless.
la croix seems like a good example.