The more interesting fact in my view is that it was trained in only.. 4 days? “Since August 28th we had been training… by september 1st we saw a step change in performance across internal benchmarks”
noma
noma’s Shortform
Some obvious takes on AI alignment that we may forget sometimes. I will just write out obvious things that are very obvious.
Let’s first write out (have chatgpt write for me) a simple impossibility theorem for AI alignment.
This is obvious
However, it becomes immediately clear upon seeing the 2nd assumption, it seems naturally that instead of trying to align any single one AI, one should be thinking about mechanism design; the same way we align humans with each other that is. How do we align humans with each other? reputation, law, and repeated interaction, treaties, deterrence, verification, trade dependence, institutions, competition, defectors, whistleblowers etc.
E.g., you can imagine that you train one set of AI’s to be defectors while another set of AI’s are trained to maximally pursue utility. There may be AI’s operating as “States” and AI’s that behave like “Cops”, etc.
If there are people working on agent foundations, then the problem of aligning agents with each other fundamentally mathematically equivalent between AI’s and AI’s and AI’s and humans. This is again an obvious statement. If you take seriously the realms of mathematics and the realms of decision theory, and you write out ‘here is an agent with some utility’ and then continue forwards with your reasoning, and you assign no specialness to humans or AI’s anywhere; then there is no reason to treat humans as different from AI’s.
This implies that if humans can be aligned with each other, then AI’s can be aligned with humans. Sometimes I think, it can be useful to write out the obvious. Hope it wasn’t too obvious however.
Terence Tao talking about the Navier Stokes solution and AI
https://mathstodon.xyz/@tao/117207849921390904
Feels like he has seen the proof/lean output.
Robocurve is working to be the METR of robotics progress, they show that GPT-6 Astra scored 95% on a robot control task, up from Fable 5.1′s 40%, with 6.2x fewer output tokens at 2.3x lower cost. Recent LLM releases seem to be able to control robots somewhat competently out of the box without any embodiment gap, and has moved my own timelines for a genuine industrial transformation of the economy forward
https://tweetviewer.com/twitter-viewer/chooi_jeq?tweet=2096064315115839904
I’m personally not concerned about AI until robotics starts to work. We are still nowhere with robotics, although there is progress