gwern
My main mistake in 2022 was not appreciating how LLM pretraining would affect the concepts available to an AI. Namely, by the time RL started, the systems would already know about the “reward” concept.
I’m reminded of my 2024 comment on the progression of model-free to model-based systems and thus increasing sophistication about reasoning/planning/searching and ability to make reward the optimization target: https://www.alignmentforum.org/posts/yQSmcfN4kA7rATHGK/many-arguments-for-ai-x-risk-are-wrong?commentId=4bNxEBNCnjuGwSCgi
It sounds like you could potentially fix this by simply training with Synth-ID on-policy.
Yes, that was part of what I was getting at. BNNs cannot implement weight sharing in a direct way like ANNs, in instantiating a bunch of copies in parallel in VRAM; they could do it by a recurrent approach, because what is a RNN but a very wide NN with weight-sharing unrolled sequence-wise? Except then that would require a large number of serial steps—serial steps which would blow any latency budget. So, pace Steve Brynes’s discussion of things like the callosum and where the parameters go in brains, they might have to do pseudo-weight sharing by just replicating a lot of brain regions with similar-ish parameters—and boom, their parameter count spikes massively, and the tighter the latency budget, the worse it gets because the less sharing remains possible through layer-wise or recurrent iteration.
Does this really work? It doesn’t work for non-code things; ask Sol or Fable to ‘write like gwern’, and assuming they don’t refuse outright because I am a living author, it is certainly not a gwern-like output ‘to the extent you can’t tell whether gwern wrote it’! Or if it does work, what would make code different?
Would it be possible to use something like Maia, instead of Stockfish? Then you could more plausibly get ‘higher human-strength moves’, where you are not left trying to followup impossibly sharp tactical play you can’t pull off without a chess engine, or other issues with doing direct naive behavior cloning from superhuman chess engines.
(Review moved.)
Or, Zhihu has a very low weight in the training set of LLMs.
This would be my default expectation: “Zhihu is hard to crawl in some way and so just doesn’t get into training datasets”. For example, Twitter is notably absent from most LLMs… Because Twitter invests a lot of effort into blocking crawlers in order to preserve tweets for Grok and do price-discrimination on the API. This means that people who primarily tweet will be under-represented in LLMs. (Even Grok doesn’t actually seem to train much on tweets.)
You also have to write about yourself. I expect for a lot of these ‘project only’ names, there just isn’t anything about them online—as opposed to the project. What else is the LLM going to say...? If someone translates a bunch of essays, but doesn’t say a word about themselves, what else is the LLM going to say other than talk about the people who wrote the essays?
Maybe source from https://arxiv.org/abs/1202.3936 https://gwern.net/doc/math/2013-hisano.pdf as pre-AI sets of conjectures to monitor or target? (I have many errors listed in my math error essay but not sure how useful the ad hoc set is compared to the Hisano & Sornette work.)
No, Guardian Angels. But to fix pretraining/dynamic evaluation, you need to enrich the principal’s data a lot, I think, and dumping in fulltext of references is a good way to ensure the LLM personas have access to the principal’s context and avoid encouraging confabulation. (Gwern.net essays/annotations/Wikipedia serve as a kind of implicit ‘reference wiki’, as do the analyses/writeups we are having the LLMs generate to reverse-engineer writing.)
You did say it was cheap. Most RSS feeds are just not that big, even with fulltext. (For a project, we’re taking my entire history of web writing and trying to put in fulltext of all links and context and elaborate metadata with verbose XML formatting… and it still winds up being only like 10b tokens.)
And I guess if you do it in the background, you can probably find an even cheaper LLM somewhere. (Does OA still do that ’50% off’ batch background thing?)
Also, note that if you summarize all the items, not just the ones the user is about to look at, you now have a useful resource for searching, embeddings, recommendations, etc.
1s is still annoying UI jank and lag. If it’s cheap, why not just run it all in advance and cache the result? RSS is the perfect case for this because the items don’t change.
sometimes just serve that version. I assume that’s what Gwern is doing
Correct. I serve Markdown in two ways: first, every page includes the HTML-standardized metadata, which points to the Markdown file version of pages (using a simple template line:
<link rel="alternate" type="text/markdown" href="https://gwern.net$url$.md” title=”Markdown source of ‘$title-plain$’ page”>); second, I support standard HTTP content-negotiation, so any HTTP request which header specifies anywhere in it that it would accept atext/markdownortext/plainresult will then be given the.mdcontents instead. (Done in nginx where this is unpleasantly complicated.)So while the features may be a trifle obscure to any given programmer (although not web developer), they are well-known standardized features going back decades being used as they are supposed to be used, and as such should be entirely obvious and easy to use for coding agents (eg. adding
text/markdownto an agent’s HTTP Accept header downloads is basically free and it should be able to blindly put that in every web request without breaking anything).
Most obvious problem to me is that if you just did a finetune on your own normal writings, then “Rewrite this passage in your own words:” is very out of distribution, because you’re trying to create a persona or base model. But no one ever talks like that in those settings; only assistant chatbots converse in these kinds of abusive, peremptory prompts. You never talk to yourself or write that in your writings (do you?), so whatever follows is going to be strange.
If you are trying to get a ‘rewrite this to sound like me’, you might want to try something like creating a small dataset of of rewrite examples with that exact formatting, where you create the pairs by grabbing random paragraphs from your own writings and have a chatbot rewrite them to be ‘better’, and reversing them. Now when you do the atomic example, it’s a well-understood task with many examples of what to do, and it should behave more normally, like ending after 1 translated paragraph with quotation marks and EOT.
For GA, we’re experimenting with heavy use of role tags, which let us synthesize transcripts of things like the ‘LLM agent’ researching/writing stuff and then ‘Gwern’ rewriting it or writing a final essay. Seems to be working.
(I have also heard that there may be some things wrong with the TMI infrastructure where it trains fine and will run locally fine, but their runtime deployment is screwed up subtly. I doubt this is the case; but you could try downloading the Kimi checkpoint to run locally slowly, just as a sanity check.)
‘AI allegory steganography’ in Claude short stories in the Unslop contest?
I’ll outsource my opinion on that to Nesov. I like density / fine-grained sparsity in principle, and since Western hyperscalers are so hardware-advantaged, I tend to expect less coarse-grained sparsity than Chinese labs are forced into.
Sure, why not. If you are bored and killing time, you may be quite curious—far more curious about the teacher’s (hopefully entertaining) past of drug use and rock and roll and hedonism than about the assigned curricular material… Nevertheless.
I don’t buy the intro example as being analogous to your two posts, or indeed, much of a conundrum about Bayesian epistemology.
The question was obviously bad to ask because it was either asked in bad faith or to kill time/boredom. You don’t need ‘advanced epistemology’ to note that the teacher’s personal anecdotal testimony has only epsilon bearing on the truth of DARE, because the badness of drug addiction is based on the experiences of billions of people over millennia and a vast amount of research, and this is as obvious as, say, the Big Bang. ‘Teacher, you say that the universe started in a Big Bang. But have you ever seen a Big Bang yourself?’ Such a question should not be dignified with an answer. Or Christopher Columbus, or, or.… in fact, ~100% of the things taught in middle school have no relationship to the teacher’s personal testimony. (Even in things like music or gym class.) Every middle schooler knows this and would regard as insane as a classmate who asks if the teacher had gone to the moon, and when the teacher responded they hadn’t, then took seriously the possibility that maybe the moon is made of papier-mâché because the science teacher hadn’t personally gone there and is ‘just’ relaying the reports of people like Neil Armstrong or astronomer consensus about it being made of rocks. The claims of DARE may be wrong (and in fact I think they broadly are from what I remember of it), but the teacher’s personal drug use tells one little—and that is before you get into anything that could reasonably be considered ‘advanced epistemology’ for middle schoolers (such as selection effects—eg. if drugs really did destroy peoples’ lives with high probability, a drug-user teacher probably wouldn’t be standing there teaching them).
Meanwhile, in both your posts, your personal experiences and judgments are a major part of them. You are not a neutral far-downstream reporter of major topics; even in the sense that you are writing about research papers in the LLM post, those research areas are extremely new and controversial and complex and have results rapidly changing at an almost daily rate, so you are applying a lot of your critical judgment and beliefs in which research to highlight, how to interpret them, and how to combine them along with your personal experiences and commentary. (Consider if the teacher had titled that class day lecture “How I Stopped Being Sure Drugs Are So Bad” and filled it with anecdotes about LSD use and the latest MAPS papers on clinical trials.)
I was unaware of this screening or this show at all, and was just walking by en route to having a snack before a call; and I was so nerdsniped by this that I wound up missing my call to watch the rest of the screening. Can confirm, funniest video I’ve ever seen about AI safety, and best thing I’ve watched since Pantheon. (Think Avenue Q+Silicon Valley.) The audience seemed to agree.
I’m surprised they seem to have so little funding—only $20k on Manifest? In terms of AI safety popular outreach, this seems like it could be excellent bang-for-buck. I feel like something may have gone wrong with budgeting if this cannot get a lousy $60k after years of trying, but there’s money for everything else like PDKU or Holly Elmore or ‘high school debate/AI safety’ conferences… It may not be another Pantheon, although I would note that musicals seem to do quite well these days (eg. Hamilton, Amazing Digital Circus, Hazbin Hotel), but that’s a pretty high bar when we’re talking funding on the order of “Constellation’s annual sourdough budget”.
I was happy to donate $250 tonight and may donate more.