Long-time lurker (c. 2013), recent poster. Cunningham’s law is my friend.
For my own reference: (1) “benchmarks” very broadly construed (2) token consumptions & costs (3) satcat mass flow notes
Long-time lurker (c. 2013), recent poster. Cunningham’s law is my friend.
For my own reference: (1) “benchmarks” very broadly construed (2) token consumptions & costs (3) satcat mass flow notes
Ooh nice callout, I do remember that post of Scott’s. Unfortunately I too don’t understand enough to comment. I’ll just quote Scott to motivate the quote you highlighted:
lylebot: Yes, even pure mathematicians often resort to experiment (more often than they admit!) when they want to guess at the truth or falsity of a conjecture. But complexity theory is a special branch of mathematics — one that asks about the asymptotic behavior, not of some particular algorithm or a typical algorithm, but of the best algorithm that could possibly exist. And the set of possible algorithms is both huge beyond imagination and lacking in any simple characterization. That’s why, in the past, experimental work has not been a big help to complexity theory: because the amount of computation that you’d have to do to learn anything interesting is so astronomical. My “wild & crazy” proposal is that we at least ask ourselves whether that’s still true, given both the better understanding and the vastly increased computing power that we have today.
The article didn’t specify. I’d be comfortable betting on the latter, since $200 per solution is in the ballpark of e.g. what the recent GDM agent did. I’d guess 10-100x more for the former.
I thought about mindspace being wide and deep when reading these AMA responses by sword swallower Bambi:
I hope this is not taken as a sign of disrespect, but I feel it’s important to ask: Why?
I don’t take it as a sign of disrespect at all! It spoke to me, in the way that other performance arts speak to other people. The person who taught me to eat fire told me to “listen to the yummy feelings” when it comes to trying different skills, and this one gave me a yummy feeling, although it’s quite a yucky thing to do.
… I’m far from an adrenaline junkie. I’m actually a pretty nervous person. It’s not that different from any other high risk sport or performance art. Every football player, aerialist, figure skater, etc. also knows it’s a matter of if, not when. 36 percent of ballet dancers retire due to a career ending injury, but every one of them decided it was worth it. Had I stayed in ballet or stuck with flying trapeze, I still would have the knowledge that there is an inevitable (and in the case of trapeze, potentially fatal) injury somewhere down the line, but my love for my art and the joy it brings me and audiences outweighs the risk.
… I do remember the first time a sword went all the way in and I felt it touch the bottom of my stomach. It was a crazy feeling but I also felt absolutely overwhelmed with joy because I had a lot of trouble getting the sword past a certain point up until then, and it did feel like I suddenly unlocked something.
How did you even get into it?
My funny answer is that I wasn’t a very good ballerina so I thought I’d try another art that is uncomfortable, dangerous, and has long term negative effects on the body. But in truth, swords just spoke to me for some reason. It was a matter of trying lots of things and finding one that felt right. I’ve done aerial silks, flying trapeze, fire, human blockhead, glass walking/eating, but sword swallowing just felt right.
How did you discover you could do this? And what made you try!?
It was less about discovering that I could do it and more about deciding that I wanted to and practicing until I could. Kind of like how you don’t discover that you can do a pull up, you decide you want to do one and start with push ups, negatives, dead hangs, etc. It took a lot of practice and patience. I decided I wanted to be a sword swallower, found a mentor, and got to work.
Have you ever had an internal injury due to your art?
Not yet. But it will happen eventually. It’s like being en pointe- it’s not a matter of if, it’s a matter of when, but proper training and training can help make the difference between a minor injury and a deadly one.
… does this skill help or hurt in any other aspects of your life?
I find it helpful. The biggest thing is that it has taught me a mental “off” switch. Which is hard to explain but when I swallow a sword I consciously suppress my gag reflex and I’ve found I am able to activate that same suppression in other scenarios. I am a pretty anxious person and I have severe OCD and I am actually now able to calm myself down when I’m freaking out or consciously ignore compulsions by using the same technique I use to swallow swords. I am also on medication though, so I suggest that as a technique for managing OCD before you start swallowing swords. It’s an unexpected bonus though.
do u see any greater meaning for sword swallowing? is it merely physical or a metaphor for something?
This is an interesting question that no one has ever asked. But yes, there is something almost spiritual about sword swallowing to me. It feels like something I was always meant to do. And for me, it’s also an exercise of patience, discipline, and mind over matter. I meditated a lot when I was learning. I used to think meditation was woo woo or nonsense but it helped so much and now I am a firm believer in the power of disciplining your mind. Being able to swallow a sword is like the ultimate meditation for me. It’s a reminder that I am in control of myself. Not to get too deep or personal, but I have severe OCD and it is nice to be able to remind myself that I am able to say no to even the strongest impulses/urges/reflexes. So I guess no, it’s not really a metaphor, but it means a lot to me as a practice and I don’t view it as just another trick I can do.
Claudiness = “good at agentic tasks, but bad at vision… and also bad at math”
I’d be curious about reception to this piece from the intended audience.
How would you feel about me sharing it on r/math by the way?
At times my work makes me feel like Siri from Watts’ Blindsight, particularly with frontier models.
This is what my father could not unmake. This is what I am:
I am the bridge between the bleeding edge and the dead center. I stand between the Wizard of Oz and the man behind the curtain.
I am the curtain.
I am not an entirely new breed. My roots reach back to the dawn of civilization but those precursors served a different function, a less honorable one. They only greased the wheels of social stability; they would sugarcoat unpleasant truths, or inflate imaginary bogeymen for political expedience. They were vital enough in their way. Not even the most heavily-armed police state can exert brute force on all of its citizens all of the time. Meme management is so much subtler; the rose-tinted refraction of perceived reality, the contagious fear of threatening alternatives. There have always been those tasked with the rotation of informational topologies, but throughout most of history they had little to do with increasing its clarity.
The new Millennium changed all that. We’ve surpassed ourselves now, we’re exploring terrain beyond the limits of merely human understanding. Sometimes its contours, even in conventional space, are just too intricate for our brains to track; other times its very axes extend into dimensions inconceivable to minds built to fuck and fight on some prehistoric grassland. So many things constrain us, from so many directions. The most altruistic and sustainable philosophies fail before the brute brain-stem imperative of self-interest. Subtle and elegant equations predict the behavior of the quantum world, but none can explain it. After four thousand years we can’t even prove that reality exists beyond the mind of the first-person dreamer. We have such need of intellects greater than our own.
But we’re not very good at building them. The forced matings of minds and electrons succeed and fail with equal spectacle. Our hybrids become as brilliant as savants, and as autistic. We graft people to prosthetics, make their overloaded motor strips juggle meat and machinery, and shake our heads when their fingers twitch and their tongues stutter. Computers bootstrap their own offspring, grow so wise and incomprehensible that their communiqués assume the hallmarks of dementia: unfocused and irrelevant to the barely-intelligent creatures left behind.
And when your surpassing creations find the answers you asked for, you can’t understand their analysis and you can’t verify their answers. You have to take their word on faith—
—Or you use information theory to flatten it for you, to squash the tesseract into two dimensions and the Klein bottle into three, to simplify reality and pray to whatever Gods survived the millennium that your honorable twisting of the truth hasn’t ruptured any of its load-bearing pylons. You hire people like me; the crossbred progeny of profilers and proof assistants and information theorists.
In formal settings you’d call me Synthesist. On the street you call me jargonaut or poppy. If you’re one of those savants whose hard-won truths are being bastardized and lobotomized for powerful know-nothings interested only in market share, you might call me a mole or a chaperone.
If you’re Isaac Szpindel you’d call me commissar, and while the jibe would be a friendly one, it would also be more than that.
I’ve never convinced myself that we made the right choice. I can cite the usual justifications in my sleep, talk endlessly about the rotational topology of information and the irrelevance of semantic comprehension. But after all the words, I’m still not sure. I don’t know if anyone else is, either. Maybe it’s just some grand consensual con, marks and players all in league. We won’t admit that our creations are beyond us; they may speak in tongues, but our priests can read those signs. Gods leave their algorithms carved into the mountainside but it’s just li’l ol’ me bringing the tablets down to the masses, and I don’t threaten anyone.
Maybe the Singularity happened years ago. We just don’t want to admit we were left behind.
OpenAI’s Astra, their “next major model”, made progress on 10 open problems in math/TCS at a total of “roughly $2,000 at Sol API rates”. See also the 62-page stylised narration of the CoTs (not full CoTs). Jotting it down here for my own reference as part of the ongoing industrialisation of pure math research.
[These problems] have been open and have seen no progress on the main result for at least a decade, and in most cases much longer… All of these problems are of substantial interest to their respective mathematical communities, and several are of broad interest across mathematics as a whole. …
We provide new results for the following problems. …
High-dimensional sphere packing. New upper bounds on sphere-packing density down to the Cohn–Elkies threshold.
Binary and spherical codes: Exponentially improved bounds on the maximum size of binary codes at any prescribed minimum distance, with analogous results for high-dimensional spherical codes.
Non-sofic groups. A construction establishing the existence of non-sofic groups, addressing a central open question in group theory.
Connes’s rigidity conjecture. Disproof of a longstanding conjecture that certain groups are uniquely determined by their von Neumann algebras
Arithmetic circuit complexity. New lower bounds for computing the permanent using arithmetic circuits and formulas, including an arithmetic-formula lower bound of order n4/log n.
Quantum parallel repetition. An exponential parallel repetition theorem for general two-player quantum games, extending a foundational principle from classical complexity theory.
Closest vector problem. Polynomial-factor hardness of approximation for the closest vector problem, a foundational lattice question related to post-quantum cryptography.
Ehrhart’s volume conjecture. Determining, in every dimension, the maximum possible volume of a convex body whose centroid is its only interior lattice point
Multicolor Ramsey numbers. A superexponential lower bound for multicolor triangle Ramsey numbers, resolving Erdős problem 183.
Extremal number conjectures. Results on the compactness and degeneracy conjectures in extremal graph theory, resolving Erdős problems 146 and 180.
Disproving the Dinitz-Garg-Goemans conjecture with GPT 5.6 Pro, a story in four acts:


In contrast to Kerger’s 10-page effortprompt, Dmitry Rybin basically just said “do a breakthrough”.
I know counterexamples to old conjectures are becoming a meme at this point. But I really cared about this problem and spent many weeks thinking about it a while ago (in both directions, proof and disproof).
I think almost all graph flows experts thought about this problem.Ben Stephens: in hindsight, could this have been found by bruteforce? and if so, how many years ago, given feasible compute of that era?
Dmitry: Not really.
I tried brute force search myself and with old LLMs too (o1/o3). The problem is there are many degrees of freedom: costs, flows, demands to nodes. so even tiny graphs have enormous number of combinations of these params
Jacob Tsimerman, who’ll likely win the Fields Medal this year, thinks that AI “is boosting his productivity by a factor of two,” mostly by speeding up the boring parts.
Q: How does AI change math, the process or the feel of it?
A: Mostly it speeds up the boring parts. There’s the act of doing math, the professional endeavor, and the act of doing math as a fun endeavor. And I think AI helps the professional endeavor, and it gets you to the fun quicker.
… one of the primary ways that I use AI is to ask sort of dumb questions, or basic questions in fields in which I am not an expert. For example, I do a lot of research on Hodge theory. In a sense, I am an expert in Hodge theory because I’ve spent years thinking about it now. But there are still so many basics that I have to look up every single time, because I forget how the technical details work. In this way, I spend a lot of time during my research asking dumb questions, getting oriented.
Another example: I know how algebra works decently well. I’m used to working with rings. I’m used to working with schemes. I have some intuition there. Say my research needs me to work, as it did in the past, with -$p$adic rigid varieties or some analogous category, with continuous functions and their spectra, or whatever it is. Math has a lot of things that are kind of similar. And then I have intuition about how I’d hope different things might work in similar ways. Before Google, I’d have to go find experts, and they’d have to make time for me — I’d ask them my questions, and then go back and forth with them. With Google this was streamlined; I could look up books, search through them, get the basic theorems, try to put things together, learn the subject.
And now with LLMs, I just ask: “Hey, there’s this theorem in algebra. Does it basically work the same way in this other setting?”
Or: “Hey, if I have a group, and it’s this size, and it has such and such a property, does it also have this other property?”
You can plug that sort of thing into ChatGPT or Claude and it’s reasonably likely to give you something useful. And it can tell me the answer, and it can explain the answer. It can be wrong sometimes. There is a skill, a skill tree, in using it and figuring out when it’s bullshitting and when it’s wrong. But it’s immensely useful.
Q: With dumb questions as the starting point, where do you go next?
A: To break down the process, I would say there are a few different aspects of my math research workflow where AI comes in.
First, there’s the finding your bearing stage, getting oriented, where you’re figuring out what’s going on in your problem, or in your theory, whatever you’re grappling with. One part of that is determining what’s hard, what’s easy, what’s known, what’s not known. So, I’ll literally prompt the AI with: “Hey, I want to solve this kind of question. Give me an overview of what’s known and what’s hard.” And it will do that really well, consistently well.
And then I ask it more targeted questions, such as: “I was thinking of these special cases, or these kinds of analogs, which of these are known, and give me references.”
So that already is a huge time saver, partly because I can do it at scale. It’s much faster than asking a person, an expert; that would take, an email, waiting for a reply, or a meeting. Now I can try it a few different times immediately with the AI.
Second, there’s searching out a strategy, trying various types of arguments to see if they’re even in the realm of making sense. I can spell out my technique and ask if that’s been done before. I can ask, “What are the types of techniques people use?”
A third way it comes in handy is in looking up relevant material, and getting references, providing links — it’s gotten much better at providing links, but sometimes it will link to a website that doesn’t work anymore. I’ll prompt it with: “I want a result like this — is anything like this known? Please point me to references.”
And then fourth, there’s what most people call doing research. The fun parts of it, the real parts. You’ve spent months getting ready, you’re uploaded, you know what’s going on. There’s nothing left to look up. Now you have to think, and you have to come up with the right math. This is where flashes of brilliance happen, or just regular math work, whatever you want to call it. But all the fun parts of math are done here, where you’re just engaged with the problem. When I say my productivity doubles, it’s with everything leading up to this fun step.
Also:
Q: Do you think there will still be a human role in terms of ideas and creativity?
A: I think there will be a point where AI will be strictly better than humans at all aspects of math: learning, proving, coming up with the problems, aesthetics. It will just be better at everything. And this will come pretty soon,
Q: How soon is pretty soon?
A: Five years? I think in two years it might already be better than us at proving stuff. We’ll be able to say to it: “Here is a statement, go prove it.” But I have wide bars of uncertainty around this stuff. It’s hard to predict the future.
Curious what you found out?
I’m assuming your project is a continuation of this?
LW regularly surfaces years-old posts in my “For You” feed, most of which I haven’t read, with a surprisingly high hit rate of “I’d like to check this out”, e.g. ten seconds of scrolling down shows John Wentworth’s 7 year old The Missing Math of Map-Making, a few more seconds of scrolling shows DirectedEvolution’s 5 year old Build Your Number Sense, etc. Kudos to the LW team and I wish other forums did this.
Philip Kerger, who teaches industrial engineering and operations research (IEOR) at UC Berkeley, just used GPT 5.6 Sol Pro in chat(!) to one-shot prove a problem in convex optimisation that had stumped him after sporadically working on it for a year or so, and that other researchers had tried and failed at (“only last year at a the ICCOPT conference I heard someone say “we have no idea” how to solve this”). Here’s the chat producing the initial proof after 148 minutes, follow-up chat refining it after 230 minutes, GitHub with Lean code and more, etc.
Kerger’s problem description from reddit
The problem concerns deterministic zeroth-order convex optimization: Let B_d be the Euclidean unit ball in ℝᵈ, and consider all convex, 1-Lipschitz functions f: B_d → ℝ. An algorithm may query any point x ∈ B_d, and receives only the exact real number f(x), no other information (but the algorithm “knows” that f is convex and Lipschitz). The algorithm is otherwise completely unrestricted, and can use unlimited computation and memory. These function-value-only problems arise naturally when an objective is evaluated through a physical experiment or simulator. One can imagine choosing d engineering parameters and observing only the cost returned by the simulation. If evaluations are expensive (think of measuring a physical system), the natural question is how many are fundamentally required. This is formalized as oracle complexity. Specifically, this is the oracle complexity of convex optimization under an exact function value oracle.
Let Q(d, ε) denote the worst-case number of queries required to find an ε-optimal point of f. An algorithm due to Protasov from 1996 shows that order d² function evaluations are sufficient, which gives Q(d, ε) = O(d²), an upper bound on the complexity. Lower bounds were practically nonexistent for this setting, and the strongest previously applicable bound was only Ω(d), inherited from the stronger first-order oracle model (where the algorithm receives both function values and gradients). That means we didn’t know for certain whether gradients actually help in optimization, since the function-value only and first-order oracle models have had this same lower bound, and so there was a linear gap in d in the complexity of this fairly fundamental convex optimization setting since 1996. So, can you find an algoritm that is better than Prosatov’s, and only needs d evaluations? Or can you show that no such algorithm can exist, and we can sleep well at night knowing that Protasov’s algorithm using d² evaluations is best possible? What 5.6 Sol proved is the latter.
What caught my eye was Kerger’s 10-page(!) effortprompt, available in Appendix A.1 of his paper, inspired by but significantly longer than OpenAI’s own effortprompt behind their proof of the cycle double cover conjecture:
Snippets of his effortprompt
Page 4 -- vibes-wise this reminded me of Eliezer’s old post referencing The Open-Source Wish Project, Wish For Immortality 1.1
Final page:
How did he come up with such a prompt? With Sol’s help of course
Kerger (paragraphs mine):
I basically had 5.6 Sol synthesize existing closely related work and their approaches, the past ideas I had, with OpenAI’s prompt that had a lot of the presumably important mechanisms for how exactly the agent should act. Especially from the “results that do not count” section onwards is a lot of input from Sol.
A major reason for the added length though is also the nature of the problem; Just the description of the problem and pointing to a couple existing results already put me at a couple pages. With this kind of complexity result what exactly fulfills the specifications of what is needed can be nuanced (so later there is also a lot more of “what doesn’t count”, and there are two major components to any proof namely the function class to use and the adversarial oracle strategy, about which you then need to prove things with convex geometry machinery, so that all adds to the length).
Also, this gap could have in principle been closed from either direction, which also adds a bit (though you can probably read between the lines that I was trying to push much more for a lower-bound, since an algorithm matching order d complexity would have been extremely surprising, since that would mean you can optimize without gradients just as fast as you optimize with them, which would be shocking given that so much practical optimization happens with gradients)
5.6 Sol was a lot better than 5.4 and 5.5 at this, although the effortprompt probably helped a lot
when I was working with 5.4 and 5.5., it kept returning to me with “Here’s great progress! The only step left is to prove this one lemma here, and here’s why proving that lemma would get the result”. Then I’d dive into trying to prove that lemma together with 5.5 and it would go nowhere. I’ll share here a chat with you that I had with 5.5 Pro Extended, where I actually told it to use max of affine functions as the hard function class (which is what ended up working and is in the preprint): https://chatgpt.com/share/6a592503-2d60-83ea-8f00-9ed96b331b16
it thought for a whole 4 minutes in its initial response. But, my prompting was also missing a lot of the what counts as a solution, only return when you have found definiteve answers, explore lemmas like this, and so on. Disclaimer, I was a bit frustrated at this point in time and was just trying to get the model to dive deeper on a couple of different approaches in sessions I was running in parallel, so most of my replies there are basically just “keep going”, especially since I had already told it what I want it to do in the initial prompt.
It cost him 10% of his $200 Pro plan (with the initial 5h/day cap)
Kerger (paragraphs mine):
Between $20 to $200 depending for me depending on how you want to look at it. I got the $200/month subscription specifically to try tackling some problems together with AI.
This specific project used under 10% of my weekly limit on there (though OpenAI messed with limits, increasing them, during this project, so I’m not sure if that’s 100% accurate). Total time was about 2.5h 5.6 Sol Pro use on the initial proof outline, 4h 5.6 Sol Pro on an improvement on the accuracy requirement, ~2h use of verifications and proof audits, ~4h use assisting for the Lean development (these later ones are estimated, and only the time of the actual active tool use, whereas the total time including my own work was of course much more). So something on the order of no more than 15h of actual time.
Even with the initial 5h cap per day they had (and then lifted), that uses 3 days = 10% of the monthly pro plan, so from the user side I can in some way say about $20 (though probably OpenAI’s compute cost much more than that!), or the $200 since I did get the pro plan for this.
As a point of reference, it cost Google DeepMind’s Gemini 3.1 Pro-based agent a few hundred dollars per proof of the 9 open Erdős problems it autonomously solved out of the full set of 353 in the open-source Formal Conjectures repo, with predictably very large variance:
As before, Sol’s proof didn’t really use nor create new techniques:
Lastly, some important comments about the work relating to AI capabilities: In a lot of cases, proving lower bounds like this result relies on finding that right construction that works (in this case, family of difficult functions and a strategy for how an “adversarial” oracle should answer queries from an algorithm to reveal minimal information) and then proving things about it. There are only so many function classes which would be reasonable to look at (here, quadratics for example would have also been reasonable with order d² degrees of freedom, or any variation of maxes of some simpler families of convex functions as well), but the actual proof mechanics once the “correct” function class and correct strategy for adversarial oracle answers is found are often not so complicated, and often employ existing results from convex geometry or similar (this is also the structure of two previous but much more niche, less important results of mine). So I wouldn’t really say that this result is using or creating some fundamentally new techniques in convex geometry or optimization theory.
What this means from my perspective is that if a result is attainable with existing techniques, modern AI methods will be able to solve those problems. I don’t think researchers in math/TCS will be made obsolete, but I think it will instead no longer make sense to work on any low-hanging, or even medium-hanging (you know what I mean) fruit. We’ll be needed for problems where actual novel approaches are needed.
But frankly I don’t think this is needed for even Fields-level contributions. To quote my own quick take:
Think about what a (static!) automated Jean Bourgain-toolkit interpolator could do
What’s the most impressive research-y feat interpolating AIs can theoretically do, fixing their training data to (say) today?
I don’t have a good sense of this in general, but in pure math it’s probably at least Fields medal-tier, if laudatios like that of Akshay Venkatesh are anything to go by:
Akshay Venkatesh stands out for the startlingly original way he has connected number theory problems to deep results in other areas. Far from using them as “black boxes” to crank out solutions, Venkatesh brings fresh insights to the results and highlights their unexpected connections to number theory. In this way he has made striking advances in number theory while also greatly enriching other branches of mathematics.
This, and the rest of the laudatio, reads like a very souped-up version of what OpenAI’s recent internal model did with Erdos problem #90 (itself a souped-up version of what GPT-5.4 Pro did with problem #1196), give or take a few Gowers-hints. One of Venkatesh’s advantages over other top-tier mathematicians is his sheer range, the thing frontier models do vastly better than humans at.
And I can imagine, for instance, 2 years of advancements enabling a frontier model to skillfully deploy the late legendary Jean Bourgain’s toolkit. Bourgain was regularly spoken of by other world-leading mathematicians as “effectively a god”. Terry Tao:
When I was a graduate student in Princeton, Tom Wolff came and gave a course on recent progress on the restriction and Kakeya conjectures, starting from the breakthrough work of Jean Bourgain in a now famous 1991 paper in Geom. Func. Anal.. I struggled with that paper for many months; it was by far the most difficult paper I had to read as a graduate student, as Jean would focus on the most essential components of an argument, treating more secondary details (such as rigorously formalising the uncertainty principle) in very brief sentences.
Tao goes on to describe Bourgain’s style and toolkit:
I began to realise that Jean had a certain collection of tools, heuristics, and principles that he regarded as “basic”, such as dyadic decomposition and the uncertainty principle, and by working “modulo” these tools (that is, by regarding any step consisting solely of application of these tools as trivial), one could proceed much more rapidly and efficiently. By reading through Jean’s papers, I was able to add these tools to my own “basic” toolkit, which then became a fundamental starting point for much of my own research. Indeed, a large fraction of my early work could be summarised as “take one of Jean’s papers, understand the techniques used there, and try to improve upon the final results a bit”.
In time, I started looking forward to reading the latest paper of Jean. I remember being particularly impressed by his 1999 JAMS paper on global solutions of the energy-critical nonlinear Schrodinger equation for spherically symmetric data. It’s hard to describe (especially in lay terms) the experience of reading through (and finally absorbing) the sections of this paper one by one; the best analogy I can come up with would be watching an expert video game player nimbly navigate his or her way through increasingly difficult levels of some video game, with the end of each level (or section) culminating in a fight with a huge “boss” that was eventually dispatched using an array of special weapons that the player happened to have at hand.
Imagine unleashing thousands of Bourgain-toolkit interpolators math-wide in 2028. I think I’m being conservative here, not assuming continual learning or whatever, not even assuming anyone else’s toolkit, and yet I still find it hard to imagine how transformative this would be. And this is just for pure math.
Look. I’m the last person who’s going to deny that the road we’re on is littered with the skulls of the people who tried to do this before us. But we’ve noticed the skulls. We’ve looked at the creepy skull pyramids and thought “huh, better try to do the opposite of what those guys did”. Just as the best doctors are humbled by the history of murderous blood-letting, the best leftists are humbled by the history of Soviet authoritarianism, and the best generals are humbled by the history of Vietnam and Iraq and Libya and all the others – in exactly this way, the rationalist movement hasn’t missed the concerns that everybody who thinks of the idea of a “rationalist movement” for five seconds has come up with. If you have this sort of concern, and you want to accuse us of it, please do a quick Google search to make sure that everybody hasn’t been condemning it and promising not to do it since the beginning.
We’re almost certainly still making horrendous mistakes that people thirty years from now will rightly criticize us for. But they’re new mistakes. They’re original and exciting mistakes which are not the same mistakes everybody who hears the word “rational” immediately knows to check for and try to avoid. Or at worst, they’re the sort of Hofstadter’s Law-esque mistakes that are impossible to avoid by knowing about and compensating for them.
And I hope that maybe having a community dedicated to carefully checking its own thought processes and trying to minimize error in every way possible will make us have slightly fewer horrendous mistakes than people who don’t do that. I hope that constant vigilance has given us at least a tiny bit of a leg up, in the determining-what-is-true field, compared to people who think this is unnecessary and truth-seeking is a waste of time.
David Reinstein’s tool might be of interest to readers https://uj-ai-wealth-philanthropy-steelman.netlify.app/
This is getting very tangential, but the moment I saw “Xiaoguang Li” I thought of the thousand year old vampire quote:
I once lent Xiaoguang “Mike” Li my copy of “Probability Theory: The Logic of Science”. Mike Li read some of it, and then came back and said:
“Wow… it’s like Jaynes is a thousand-year-old vampire.”
Then Mike said, “No, wait, let me explain that—” and I said, “No, I know exactly what you mean.” It’s a convention in fantasy literature that the older a vampire gets, the more powerful they become.
Yeah, to corroborate your 1st bullet point, I came from the private sector where I spent over a half decade as a data analyst worrying about and ensuring trustworthiness of numbers reported to executive teams making lots of tight-feedback loop high-stakes decisions, and when I pivoted to my current academia-adjacent career path I was shocked to see how much worse data trustworthiness was in comparison even for supposedly high-quality journal articles when I started digging into things.
Byrnes made a related point I think (more here):