My speculation and guess that Zeke is Ezekiel Reffe-Hogan—of AI governance project. Hunch based on their association with Kelsey Piper, generally associated with AI stuff etc
winstonBosan
Gwern has retired from fulltime writing + pseudonymity and is now working full time on Guardian Angel—definitely didn’t have this on my bingo card.
Neat! I’d love to see Hamburg and Lubeck return to their former glories—but there is a wrinkle comparing China of 1970/1980s to Europe of today.
China in 1980s, Singapore in the 1960s, and to a lesser extend Dubai in 90s, is that they are desperate in a way that the Europeans are not. And they have taken advantages of having either a vast pool ofexploitablecheap labor in its hinterland, or exceedingly liberal immigration policies (Will the Europeans of today want this?). SEZs cannot be cutoff from the mother country/host country, and I suspect Europeans are just not desperate enough to make a gamble like this.
Continued finetuning and RLVR and distillation might be one answer—the models do get progressively better and cheaper, but external parties no longer get the largest model built on the most recent pretraining run. Also, I don’t think the latter is that far off.
Calling a whispering earring by another name does not make it any less disempowering. i share the worries about this kind of self inflicted disempowerment.
As someone who has both had a professional career in and out of China, I am utterly confused by the statement that the China today thinks ”rest of the world [is] low in quality”.
Care to share why you think that is the case?
I am sympathetic to the revolutionary vanguard approach. But MAPLE seems to have even less success in the approach than the SRs of the russian revolution. At least the SRs tried going to the people first—telling them “peasants, you are being exploited and you could change your condition!”.
From the outside, I don’t see any saintly accomplishments that look like “trying to apprehend alignment” at a robust mechanical level.
I think read the docs is a fine advice. But docs rot, and most docs are bad. If possible, read the code?
Reading the code doesn’t mean grokking the entire cpp code base while reading by candlelight in 1000-day monk mode. Reading the code means understanding the code path—descending from your first contact point down to the stateless functions that does the heavy lifting. This can be done with an actual debugger and stepping through each stage and see how the sausage is made. Reading the code is slow, sometimes painful. The alternatives are worse.A bad advice is to use a “better” language. I guess that Raemon uses weakly typed languages beyond the base rate of js/python users. Many of the complaints here don’t really make sense if you write in apl, haskell or even rust. When you create the right abstraction, or best, have no abstraction, there is no meta state of the program to keep in your mind.
You probably don’t like the term LLM because it doesn’t describe capability. And most model are multimodal these days, so it is not just natural language.
You also wouldn’t like the term Autoregressive/Next-token predictor. Still because it says what it does, not what it is capable of.AI is a pretty good term. As overloaded as it is.
go-away is my personal choice.
Doesn’t require weird js and text mode browsing like Anubis. Widely(ish) used. Not nuclear like anubis.
Not a downvoter, but I am put off by things like:
| Runs 93x faster than Zephyr 7B
On a…. What? A potato? A consumer gpu that doesn’t fit all of the 7B model so it is mem-moribund? Things with “patent pending” (nothing wrong with patents!) and permitting grad students to use it “for their degrees”. Just enough little vibe nudges that I feel confused and unmotivated to actually read the code/paper.
FYI there is branching on Claude desktop.
Fair. In Sarah Constantin’s terminology, it seems you aspire to “potentially take a stand on the controversy, but only when a conclusion emerges from an impartial process that a priori could have come out either way”. I… really don’t know if I’d call that neutrality in the sense of the normal daily usage of neutrality. But I think it is a worthy and good goal.
I don’t think Cole is wrong.
Lesswrong is not neutral because it is built on the principle of where a walled garden ought to be defended from pests and uncharitable principles. Where politics can kill minds. Out of all possible distribution of human interactions we could have on the internet, we pick this narrow band because that’s what makes high quality interaction. It makes us well calibrated (relative to baseline). It makes us more willing to ignore status plays and disagree with our idols.
All these things I love are not neutrality. They are deliberate policies for a less wrong discourse. Lesswrong is all the better because it is not neutral. And just because neutrality is a high-status word where a impartial judge may seem to be—doesn’t mean we should lay claim to it.
The claim about “no systematic attempt at making a good [prompt]” is just not true?
See:
Is this earth shattering as an observation? No.
But it is a good retelling of a phenomenon that most of us experience personally (density of time lived corresponding to novelty) tied in with interesting personal backstory. This is a good post that I enjoyed reading.
Good suggestion. Thou it was mostly as a tribute to a programming blog that I really appreciate. I expected no one else to read this but me.
TLDR: random is roughly 477 elo. (See details and caveats below)
See link, where someone matches weird chess algos against various dilutions of stockfish (stockfish at lvl n with x% of random move mixed in). So this elo is of course, relative to various stockfish dilutions at the time of recording.
https://www.youtube.com/watch?v=DpXy041BIlA
Revolutions
I like how it is a focused podcast on just the most important 10 revolutions of human history. Told humorously and researched thoroughly, it aims to teach the general truth through indepth case studies.
It teaches history very well. By using the revolution as the narrative skeleton, you are kept on track by the chronological evolution of the event. While learning the people and systems through context and fractal expansion on the details of the said participants.Yes, mostly.
People with a vague interest in history and likes stories told well.
I didn’t vote disagree, but my concern is that, if this were to work, it is just gradual dis-empowerment by another name and something very whispering earring shaped.