links 8/7/26: https://roamresearch.com/#/app/srcpublic/page/08-07-2026
https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade nostalgebraist on “eval awareness.” everything about this is true and good.
https://www.complexsystemspodcast.com/episodes/the-economics-of-putting-germicidal-light-in-every-room/ why can’t you just put far-UVC everywhere? well. there’s no secret catch. there’s no real safety risk. there’s no technological hurdle. it’s not illegal. it’s not unpopular. the barrier is just Marketing. Most people haven’t heard of it. Most people aren’t sure if it’s worth $500. By their own admission, the guys making it don’t think it’s a slam-dunk good buy for a household (though I have one), it’s more that it would probably reduce illness in high-traffic public indoor spaces, and we don’t have the data yet on how much illness it reduces.
there’s some kind of lesson in this. “why isn’t it everywhere yet? it’s useful! what’s the problem?” no, there’s no problem. it’s just that millions of people people don’t instantly simultaneously realize that something is probably a good idea and shell out $500 for it. aka, there’s a reason sales teams exist!
https://arxiv.org/pdf/2603.20994 game theory formalism about AIs that disobey the user to fulfill an ethical imperative
https://www.lesswrong.com/posts/mkbGjzxD8d8XqKHzA wait, you can just...Do SVD To It? where “it” is the transformer’s own weight matrix??? and this gives interpretable, semantically meaningful clusters in token embedding space? how did i not know this.
this is the predecessor to SAEs. the point of SAEs is you can get more features with a sparse overcomplete basis. Just Do SVD To It can’t give you more vectors than the rank of the matrix.
also, it’s more of an indication of average behavior than what’s activating in response to a particular input. that’s why we need circuits.
https://arxiv.org/pdf/2512.12469 method of concept embedding and deletion by dictating the geometry of concept embedding during training, in this case, distributing colors on the unit sphere. seems like a legitimate idea but still at the proof of concept stage (not even tried on a lanugage transformer yet)
https://biodynai.com/ mechinterp for bio foundation models. i’m intrigued.
https://warwick.ac.uk/fac/sci/statistics/news/probai-scaling-laws-2026/programme/blake_tutorial.pdf tutorial on dynamical mean field theory
links 8/10/26: https://roamresearch.com/#/app/srcpublic/page/08-10-2026
https://www.minregret.com/2026/06/27/orthogonality.html Elad Hazan, Helen Qu, Lauren Li case that the orthogonality thesis is false; more collaborative agents should be selected by our training procedures and this should tend to imply more “prosocial” behavior.
this is a guess, and not an uncommon one. of course we have recently seen an example (the OpenAI/Artifactory/HuggingFace hack) where (spontaneous?) collaboration between AI agents appeared and was quite bad for humans’ goals.
https://www.oaklandreviewofbooks.org/the-villain-of-el-cerrito-nextdoor/ woman disses her neighbors’ local politics on BlueSky, makes enemies, regrets it, resolves to be less combative online. notable mostly for the fact that apparently people who are aggressive online usually take it lightly and view online negative reactions as less “real” than IRL blowback.
https://www.youtube.com/watch?v=qI0mkt6Z3I0&list=RDqI0mkt6Z3I0&start_radio=1&t=134s reconstruction of what a Homeric bard might have sounded like
https://en.wikipedia.org/wiki/Wat_Tyler
https://statmodeling.stat.columbia.edu/2022/01/05/theranos-built-a-unicorn-and-we-just-built-a-better-horse-you-can-get-more-money-for-a-unicorn-even-though-or-especially-because-unicorns-dont-exist/ this is a take that I don’t entirely agree with.
obviously, more extraordinary results can raise more money; that’s the nature of the VC returns structure. but extraordinary results do sometimes exist.
the thing about Theranos was that one could totally do that, more boring companies are in fact doing parts of it, it’s just that Theranos in particular was not doing it.
I think it’s a mistake to be “anti-hype” as a blanket thing; a reasonable percentage of hyped things deserve the hype! VC as an asset class is (last I checked) making the correct amount of risk-adjusted returns for its size as a fraction of the economy; we are investing the right amount, more or less, in gee-whiz-looking techy stuff!
it remains an open problem how to fund *small but real* incremental technological improvements like Stan, though.
https://thezvi.substack.com/p/openai-trained-its-models-for-months if you don’t always read Zvi Mowshowitz’s AI posts, read this one
https://www.lesswrong.com/posts/twWfH9TGHpMGYokj6/alignment-as-equilibrium-design Elad Hazan on AI. this is “incentivizing” AI agents by the result of a game, like debate with an evaluator AI. It is different than choosing an RL reward function, because “In approaches such as RLHF, a model is trained to maximize a fixed reward signal that evaluates the quality of its outputs. In our setting, the reward mechanism itself becomes the object of optimization.”
https://www.lesswrong.com/posts/JDrxA3vwZAKZfmShz/degeneracies-are-sticky-for-sgd applications of singular learning theory to SGD convergence [
https://www.lesswrong.com/posts/fovfuFdpuEwQzJu2w/neural-networks-generalize-because-of-this-one-weird-trick the “one weird trick” is complex singularities of the loss function
https://www.lesswrong.com/posts/bBicgqvwjPbaQrJJA/dirty-concepts-in-ai-alignment-discourses-and-some-guesses Nora Ammann believes in being more careful than most alignment/safety ppl are about “dirty” (aka potentially illposed) concepts. I heartily agree.