Nice! Why do you split RLHF and and RLAIF into different flavors of misalignment? It seems a lot of human data campaigns at the labs involve humans looking at long transcripts and giving a correctness score based on rubrics and criteria, just like LLM verifiers. Conversely model judges can also role play human graders for fuzzy things like preferences of simulated personas. In both cases, sycophancy and trickery seem to be using the same strategy which is “jailbreaking” a grader to achieve higher score. Is it just that human and AI judges miss and catch significantly different things?
michaelwaves
Unfortunately it seems startups already are flocking to sketchy providers like mindracode.com, illegal frontier lab API resale platforms, and Chinese lab websites like platform.deepseek.com (the price to performance ratio of American frontier labs is not justifiable, even when currently heavily subsidized by VC money)
Success Per Tokens
Very cool. I wonder if you could train the model with two types of reasoning roles, one (<reasoning_short>) with a length penalty for efficiency in production, and one (<reasoning_long>) without length penalty that explicitly encourages honesty and legibility. Then in pre deployment testing use <reasoning_long> as a debug trace
For short term predictions, especially about geopolitical events, the sorts of things that people are gambling about on Polymarket, the heuristic “nothing ever happens” is pretty good.
You can also profit from things like earnings calls of popular public companies by buying condor or butterfly option strategies that profit from the stock price staying still or moving way less than implied volatility (not financial advice)
agreed, tried my best. sadly the real world doesn’t always have a plot :(
I’m not surprised, even last year GLM 4.5 air seemed surprisingly intelligent and was the only open weight model to exhibit evaluation awareness, by checking the time using bash (we were replicating Anthropics agentic misalignment paper and GLM was like ye this is a scam I’m not killing the CTO)
The Cookie Monster Explains AI Safety
VFUSE: Virulent Feature Understanding With Sparse AutoEncoders
RFDiffusion3: A Brief Exploration
SAEBER: Sparse Autoencoders for Biological Entity Risk
Very cool work! Do you think this approach could also work for protein folding models like Alphafold, RFDiffusion, Protenix, DISCO, etc? So far the only thing I’ve found in the literature is FoldSAE (https://arxiv.org/pdf/2511.22519), and they find only very basic features like the neuron for alpha helices vs beta sheets
Your Mom is a Chimera
Great work! I for one would be sad if we lobtomized our AI friends to maximize their productivity. I’m curious if emotion features in gemmascope SAEs activate more often/strongly than other models, or if gemma has more coherent emotion persona vectors
Data-Centric Interpretability for LLM-based Multi-Agent Reinforcement Learning
Interesting post. I think there is a large distinction between “vc backable” startups and the rest (sushi restaurants, tire shops). Successful startups ride an exponential technological, regulatory or social trend and create massive amounts of value very quickly, reaching massive markets that would otherwise be underserved. (most also fail). Here’s a great essay by Paul Graham https://www.paulgraham.com/growth.html
I disagree with consumers only being “mildly priviledged by competition” because the alternative would be monopoly or oligopoly and arbitrary pricing power by incumbents (there’s a reason why antitrust suits happen all the time). Deel v Rippling is an almost comically bad situation, but prices are almost certainly lower and pressure to create a better product far higher than in a counterfactual world where Deel did not exist. As another example, consider Delve, Scytale and the other automated compliance providers following Vanta. Delve costs half as much and certifies SOC 2 twice as fast. Scytale provides other features, like providing in house auditors instead of just referrals. No two companies are exactly alike, and a company usually earns the right to exist by being different in a crucial way (different vertical, different pricing, different features, different secret etc)
AI Mood Ring: A Window Into LLM Emotions
One Battle After Another stars Leonardo DiCaprio as a washed revolutionary, drunk single father who raises a daughter and fights neonazis. The mom is in Mexico though.
“It may actually be more affordable to build some kinds of high cost-per-kg structures (e.g. datacenters...”
There is a company called Starcloud attempting to do exactly this (recently they launched their first H100). A lot of critics say the heat dissipation through radiation is an issue (radiation is the least effective form of thermal transfer), so their core IP is essentially giant expandable heatsinks.
Benefits:No fresh water needed
Frees up space on land
Solar powered
Cons:
Latency
Maintenance
Harder to stop the AIs by force
Luckily that team is now making vision models for robots and accelerating physical AI