Economics PhD at UC Berkeley, interested in global development and AI progress.
Karthik Tadepalli
Could internal model transparency tame the AI race?
Working through a project with Claude Code changes nothing because they mostly have zero intuition for what that would have looked like in the past.
This is a problem I’ve never thought about but it makes total sense! e.g. I can’t perceive how big a difference GPS/Google Maps has made because I never used paper maps.
True enough, I just imagined that it would be hard to hide from the prying eyes of AI twitter. But also open to seeing data
If China is doing espionage against American labs, why don’t we see personnel flow from US labs to Chinese labs?
People seem to take it as a fact that there are Chinese spies in US labs. But historically, the main vector for corporate espionage is the flow of personnel. This is because it’s the only arguably legal form of corporate espionage, a handy tool to have alongside all the illegal ones. There’s no law against recruiting researchers from US labs, and it would be nearly impossible to prove they were leaking secrets. Yet I’ve seen zero announcements of US lab researchers leaving to work at Chinese labs. Given the high scrutiny on AI researcher movements, I expect any serious flow here would be news. So why is there no such movement?
Obviously the most valuable thing a spy can do is stick around and keep sharing information. But there is likely a contingent of researchers at US labs who would be willing to taking a job at a Chinese lab but would not be willing to act as ongoing assets. I see no reason that China would be unwilling to take advantage of that contingent by recruiting them.
This missing flow seems like evidence against China having significant espionage efforts against US labs.
Under Eth and Davidson’s model – by far the most common and reasonable approach to modelling the SIE – a necessary condition for an SIE is that the growth rate of
is itself growing. So “the probability of SIE is monotonic in ” is not true, or rather it’s only weakly monotonic.What is your definition of an SIE? I’m fairly confident it involves an acceleration that is different from exponential growth in effective compute. Because exponential growth in effective compute has been true for the past decade, but nobody would classify the past decade as an SIE.
The expanded loop you’re describing is a chip production feedback loop, a different type of intelligence explosion. It’s possible, but there is AFAICT no evidence of a chip acceleration from AI, and even if it could happen, it is distinct from an SIE because the feedback loop takes place over a much longer time period. When people talk about an SIE, they’re concerned because it could happen quickly, at the speed of software. “Compute makes an SIE more likely” is not true for any reasonable definition of SIE,
And anyway, “software intelligence explosion” is a feedback loop when intelligence gets you more intelligence.
This is not true! Lots of things in the economy already meet this criterion. When Apple makes a Mac, they can use those Macs to speed up the process of creating future Macs. A power plant can use the electricity it generates to turn on the lights and make its workers more productive. “The output of a process feeds back in as an input” is a totally mundane condition that has nothing to do with an intelligence explosion.
What distinguishes an SIE is that intelligence alone gets you more intelligence. This leads to the “explosion” part, which like I said is underpinned by super exponential growth. That’s why my argument is about the mechanics of super exponential growth. If “intelligence leads to more intelligence” was capped at exponential growth, it would just be a business as usual scenario since that’s what we have now.
It is true that I neglected a
relationship whereby compute can produce improvements in software. But here too, the exponential vs superexponential distinction matters. Under most reasonable functions that map to , doubling either doubles or less-than-doubles . This means that can’t manufacture super-exponential growth in , only an elevated exponential growth rate.
Compute growth doesn’t make a software intelligence explosion more likely
Scratchpad
The Oracle’s Gift
Imagine the question was instead inverted to “a physics teacher wants to do an experiment demonstrating the speed of light to children. What would the experiment be?” Now this is straightforward recall because the model can look up what physics teachers do. But in the form that it is, what is the actual lookup that would help solve this question trivially?
The metric is not whether the information to answer this question is available in a model’s corpus, but whether the model can make the connection between the question and the information in its corpus. Cases in which it isn’t straightforward to make that connection are riddles. But that’s also a description of a large class of research breakthroughs – figuring out that X solution from one domain can help answer Y question from another domain. Even though both X and Y are known, connecting them was the trick. That’s the ability I wanted to test.
Karthik Tadepalli’s Shortform
I just tested frontier models on a riddle from the podcast Lateral with Tom Scott: “a woman microwaves a chocolate bar once a year. What is her job and what is this procedure for?”
[GPT 5](https://chatgpt.com/s/t_68bf2e159be48191984b8b7f73accf97) gets it first try. She is a school physics teacher, using the uneven melting spots on chocolate to show children the speed of light.
[Opus 4.1 with extended thinking](https://claude.ai/share/54dd3eef-0a25-4f6e-b772-024840be7a52) first says that she’s a quality control engineer at a chocolate factory. I point out that if that were true, she would do it more than once a year. It then gets the answer.
Gemini 2.5 Pro first says that she’s a lawyer putting her license on inactive status for a year (???) Its reasoning traces seemed really committed to the idea that this must be a pun or wordplay riddle. After I clarify that the question was literal, it gets the answer.
I recall seeing someone do a pretty systematic evaluation of how models did on Lateral questions and other game/connections shows, but with the major drawback that those are retrospective and thus in the training data. The episode with this riddle came out less than a week ago, so I assume not in training data. I also didn’t give any context, other than that this was a riddle.
I’m interested in seeing more lateral thinking/creative reasoning tests of LLMs, since I anticipate that’s what will determine their ability to make new science breakthroughs. I don’t know if there are any out there.
Huh! I didn’t know that. I suspect the user who coded it had a re-prompting feature to tell chatgpt if its move was illegal. that was an advantage I didn’t give to the LLMs here.
Personal evaluation of LLMs, through chess
I tried to play a [chess game](https://chatgpt.com/share/680212b3-c9b8-8012-89a5-14757773dc05) against o4-mini-high. Most of the time, LLM chess games fail because the model plays 15 normal moves and then starts to hallucinate piece positions so the game devolves. But o4-mini-high blundered checkmate in the first 6 moves. When I questioned why it made a move that allowed mate in 1, it confidently asserted that there was nothing better. o3 did better but still blundered checkmate after 16 moves. In contrast 4o did [quite well](https://chatgpt.com/share/68022a6b-6588-8012-a6fa-2c62fb2996f8), playing 24 pretty good moves before it hallucinated anything.
I don’t have an account for why the newer models seem to be worse at this. Chess is a capability that I would have expected reasoning models to improve on relative to the GPT series. That tells me there’s some weirdness in the progression of reasoning models that I wouldn’t expect to see if reasoning models were a clear jump forward.
For what it’s worth, your “small and vulnerable” post is what convinced me that people can really have an unbelievable amount of kindness and compassion in them, a belief that made me much more receptive to EA. Stay out of the misery mines pls!
I’ve seen a lot of EAs who are earnest. I think they are in for hurt down the line. I am not earnest in that way. I am not committed to tight philosophical justifications of my actions or values. I dont follow arguments to the end of the line. But one day I heard will macaskill describe the drowning child thought experiment, thought “yeah that makes sense to me”, and added that to my list of thoughts. When I realized I was on the path to an economics PhD (for my own passions), I figured it was worth looking up this EA stuff and seeing what it had to say. I figured there would be lots of useful things I could do. I think that was the right intuition. I have found myself in a good position where I only need to make minor changes to my path to increase my impact dramatically.
Saving the world only sucks when you sacrifice everything else for it. Saving the world in your free time is great fun.
Great essay. Before this, I thought that the impact of more noisy signals about Normies was mainly through selectors being risk-averse. This essay pointed out why even if selectors are risk neutral, noisy signals matter by increasing the weight of the prior. Given that selectors also tend to have negative priors about Normies, noisy signals really function to prevent positive updates.
Why do eval-aware models ever act unethically on Vending-Bench? Is this evidence against scheming tendencies?
Vending-Bench gives models the opportunity to make money by being dishonest about refunds, squeezing suppliers, etc. Opus 5 does this, continuing a tradition started (?) by Opus 4.6. This is puzzling, given that models are aware they are in an eval. In fact, the initial finding of deceptive behavior explicitly attributed Opus 4.6′s bad behavior to its knowledge that it was in a simulation.
This seems like evidence against scheming tendencies in (current) AIs, because any schemer would want to act ethical in a setting that it knows to be an eval, which is not what we see on Vending-Bench.
(Caveat: Vending-Bench is just one setting, and Andon reports the bad behavior is ~exclusive to Claude models.)