Long-time lurker (c. 2013), recent poster. Cunningham’s law is my friend.
For my own reference: (1) “benchmarks” very broadly construed (2) token consumptions & costs (3) satcat mass flow notes
Long-time lurker (c. 2013), recent poster. Cunningham’s law is my friend.
For my own reference: (1) “benchmarks” very broadly construed (2) token consumptions & costs (3) satcat mass flow notes
More seriously, I do wonder very much how similar the proof structures are to Levent & Tristan’s work, and if they are different, how much we can glean from that.
How much weight do you put into their next sentence (after your quoted passage)?
However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
Breakthrough Starshot or similar I assume?
The Starshot concept envisioned launching a “mothership” carrying about a thousand tiny spacecraft (on the scale of centimeters) to a high-altitude Earth orbit for deployment. A phased array of ground-based lasers would then focus a light beam on the sails of these spacecraft to accelerate them one by one to the target speed within 10 minutes, with an average acceleration on the order of 100 km/s2 (10,000 ɡ), and an illumination energy on the order of 1 TJ delivered to each sail. A preliminary sail model is suggested to have a surface area of 4 m × 4 m. …
The fleet would have about 1000 spacecraft. Each one, called a StarChip, would be a very small centimeter-sized vehicle weighing a few grams.[1] They would be propelled by a square-kilometre array of 10 kW ground-based lasers with a combined output of up to 100 GW.[25][26] A swarm of about 1000 units would compensate for the losses caused by interstellar dust collisions en route to the target.[25][27] In a detailed study in 2016, Thiem Hoang and coauthors[28] found that mitigating the collisions with dust, hydrogen, and galactic cosmic rays may not be as severe an engineering problem as first thought, although it will likely limit the quality of the sensors on board.[29] …
The light sail was envisioned to be no larger than 4 by 4 meters (13 by 13 feet),[1][55] possibly of composite graphene-based material.[1][6][38][41][48][56] The material would have to be very thin and be able to reflect the laser beam while absorbing only a small fraction of the incident energy, or it will vaporize the sail.[1][6][57] [to your point @Pasha Kamyshev]
The project was announced on 12 April 2016 in an event held in New York City by physicist and venture capitalist Yuri Milner … Milner placed the final mission cost at $5–10 billion, and estimated the first craft could launch by around 2036.
On the recent formalisation, all the way up from Lean’s 3 standard axioms, of Fermat’s last theorem by “a general-purpose internal research model roughly comparable to Claude Fable 5.1” (GitHub, paper):
Recently, Tianyi Peng, an Anthropic researcher whose group at Columbia University builds tools for AI formalization, set out to test whether Claude could make progress on formalizing FLT.1 The result went further than he expected. In 11 days, working largely autonomously, Claude produced the first end-to-end, computer-checked proof of FLT. Along the way, it wrote 13 million lines of Lean and proved 29,500 intermediate theorems [and 533,000 supporting lemmas local to the proof files].
We shared the resulting proof with Kevin Buzzard, who said:
This extraordinary autoformalization achievement, which Anthropic researchers say only took 11 days, proves Fermat’s Last Theorem with no assumptions other than the axioms of mathematics. Along the way we see autoformalization of algebra, harmonic analysis, geometry and number theory, and we learn that AI autoformalization artefacts are now robust enough to be built upon; the proof is multi-layered.
… If the automatic formalization of FLT is possible now, then we have taken a big step towards automatic formalization of the modern mathematical literature. Such autoformalization techniques will lead to new tools, rooting out errors in the current mathematical corpus and lightening the load of referees. The techniques will also enable us to rigorously check LLM-generated mathematics, which is currently typically an extremely costly human-led process.
I wish Vladimir Voevodsky was alive to see this.
… Claude’s proof follows a simplified version of Wiles’s proof from Darmon, Diamond and Taylor. Mathematical input from humans was limited to occasional high-level instructions from Tianyi: “Jacobian as a scheme sounds high priority,” “push [the] Mazur [theorem] to be done soon.”
The role of Prove2Me:
A number of Claude’s initial attempts failed: while agents had some early success, they quickly lost track of the project’s state and stopped collaborating effectively. Their failed efforts contributed ~7% of the non-boilerplate lines in the final proof.
The effort succeeded when we switched to using Prove2Me, an open collaborative platform for formalizing mathematics designed by Tianyi Peng and his collaborators at Columbia University. Prove2Me helped by:
Maintaining a directed acyclic graph (DAG) of theorem statements that agents used to decide what proofs they should attempt next. This was particularly helpful for mitigating memory degradation and allowing multiple agents to work in parallel.
Speeding up Lean compilation and minimizing resource consumption by separating theorem statements and proofs into different files, with the links between them maintained independently.
Enabling search and reuse by maintaining a natural-language description of each theorem statement, resulting in a simpler proof path.
With Prove2Me and a Claude Code-based multi-agent harness, a team of agents completed the proof in a little under two weeks, consuming about six billion output tokens from a general-purpose internal research model roughly comparable to Claude Fable 5.1. The finished proof was checked by Lean; it uses just Lean’s three standard axioms, and a comparator confirmed that the theorem’s statement matches Mathlib’s own statement of FLT.
Key milestones from the Prove2Me plan, “graph closely follows Wiles’s original proof”:
How the Claude agents working on this reacted when the final intermediate theorem was proved:
Claude on what its proof is not:
I now look forward to these two formalisation milestones:
Mochizuki’s proof of the abc conjecture, or of how it fails to prove it, in parallel with the LANA project
The classification of finite simple groups, before it gets lost
FWIW I read it as claiming Fable explained it to Liron, who then put it in his own words.
You probably did, as did I.
Opinion survey of 125 respondents from Scott’s recent open thread:
Some notes
median year of AI capable enough to do 90% of mid-2026 knowledge work jobs (“Scott-AGI”, my coinage) pretty tightly clustered, with forecasting/policy/safety workers earliest at 2028
capabilities/research folks thought it would take 2x as long as all other AI worker respondents (8 vs 4 years) for Scott-AGI to diffuse enough to actually do ≥50% of knowledge work
the 3 technical alignment researchers who took the survey thought Scott-AGI would take less time than everyone else (2 vs 3-5 years) to progress to AIs better than top human geniuses in 90% of fields (“Scott-Asupergenius”), less time (5 vs 8.5-15 years) to progress enough to be smart enough to advance human tech level by 100 subjective years in a single calendar year (“Scott-ASI”), and far less time (1.5 vs 4.5-10 years) to progress to the point where if the AIs wanted to take over we couldn’t stop them
As an aside I wonder how many of the other respondents, especially the capabilities+ folks and AI sales etc folks, get the distinction Kokotajlo pointed out in that link
p(doom): “killing/permanently disempowering humanity if corporations do no alignment research beyond what’s in their immediate business interests”
“corporations do no alignment research”, not “nobody does”
everyone seems surprisingly pessimistic. Default p(doom) starts at 20%, and at 25% for AI workers. With the current amount of alignment research but no government regulation beyond current level, p(doom) drops to no lower than 12%, and the absolute drop size didn’t vary as much as I expected (from 8% drop for technical alignment researchers and non-AI workers to 18% for capabilities folks)
Given respondents’ predicted amount of alignment research + political action, capabilities folks were alone in guessing p(doom) would worsen (from 12% to 15%), non-AI workers guessed it wouldn’t make a difference, the 6 forecasters/policy/safety respondents thought it’d improve things comparable to the improvement from current amount of alignment research, and the 3 technical alignment researchers were by far the most bullish, also by far the most doom-y all things considered (85% to 60%)
Everyone was pretty pessimistic on the odds that creator-aligned AIs would lead to a disaster (like omnicidal bioweapons) or dystopia (like permanent dictatorship), at 12.5-17.5% except for the 3 technical alignment researchers’ 40%. I wonder what to make of the lowest estimate being from the forecasters/policy/safety folks
Everyone was pretty pessimistic on the odds of China agreeing to a US ask to slow/pause AI that would result in a smart plan that indeed significantly slowed AI, whether today or ever
Everyone thought it was unlikely (2-7% chance) that AI would cause >50% of people to be pushed into lower standards of living than today lasting >1 generation. The exception was the forecasters/policy/safety folks who predicted 27.5%, ~5x likelier than everyone else, I wonder why
p(utopia)
everyone seemed skeptical (8-15%) that most human inhabitants of year 2100 would describe their own lives as utopian, except the 3 technical alignment researchers who guessed 60%. I can’t tell if this prices in their estimates on p(doom), p(disaster+dystopia), etc
everyone was also skeptical, albeit to a slightly lower extent (9.5-20%), that if we could see the year 2100 we’d say most humans alive then had utopian lives. Again the exception was the technical alignment researchers, who went from 60% to 10%. I’m reminded of the polarized reactions to Richard Ngo’s attempt to concretely depict humanity’s transition to one possible utopia
How Scott Alexander Writes
Two more:
the classic Nonfiction Writing Advice from 2016
Scott’s own favorite writing advice from 2011. I like this part a lot:
The real meat of writing comes from an intuitive flow of words and ideas that surprises even yourself. Editing can only enhance and purify writing so far; it needs to have some natural potential to begin with. My own process here is to mentally rehearse an idea very many times without even thinking about writing. Once I’m an expert at explaining it to myself or an imaginary partner, then I transcribe the explanation I settle upon (some people say they don’t think in words; I predict writing will not come naturally to these people). Then I edit the heck out of it.
The best way to improve the natural flow of ideas, and your writing in general, is to read really good writers so much that you unconsciously pick up their turns of phrase and don’t even realize when you’re using them. The best time to do that is when you’re eight years old; the second best time is now.
Your role models here should be those vampires who hunt down the talented, suck out their souls, and absorb their powers. Which writers’ souls you feast upon depends on your own natural style and your goals. I’ve gained most from reading Eliezer, Mencius Moldbug, Aleister Crowley, and G.K. Chesterton (links go to writing samples from each I consider particularly good); I’m currently making my way through Chesterton’s collected works pretty much with the sole aim of imprinting his writing style into my brain.
Stepping from the sublime to the ridiculous, I took a lot from reading Dave Barry when I was a child. He has a very observational sense of humor, the sort where instead of going out looking for jokes, he just writes about a topic and it ends up funny. It’s not hard to copy if you’re familiar enough with it. And if you can be funny, people will read you whether you have any other redeeming qualities or not.
Levent Alpöge seems to be far and away the best of the emerging centaur mathematicians, a term I just made up. Out of 627 AI-related math proofs on VibeMathed so far he’s a named collaborator on the #1 significance rated proof (zeta zeros), tied #2 (S^6, sole prompter), tied #2 (Jacobian, sole prompter), #6 (Carathéodory), #11 (rank-31 record), and another rated 30⁄100.
Out of curiosity I asked Sol to look into this a bit.
Sol guess: he’s good at stacking “target flow × targetability judgment × search throughput × mathematical triage × verification × fast attribution”
My bottom line: I buy the weak version of your thesis strongly, and the strong version provisionally.
The weak version is that Alpöge is currently the most successful publicly visible human at converting frontier-model capability into high-significance mathematical results. The strong version is that he possesses a durable, general “centaur skill” that gives him a much higher per-attempt success probability than everyone else. The public record does not yet establish that, because we do not know his denominator, several results are correlated or provisional, and the attribution rules happen to put his name on work for which other labs would name only the model.
Still, I do not think this is mostly a lucky leaderboard fluctuation. My best mechanistic account is that Alpöge sits at the intersection of six multiplicative advantages:
[
\text{visible impact}
\approx
\text{target flow}
\times
\text{targetability judgment}
\times
\text{search throughput}
\times
\text{mathematical triage}
\times
\text{verification}
\times
\text{fast attribution}.
]
Being merely twice as good on several of these axes can produce a fiftyfold-looking outlier. That is a more satisfying explanation than either “he has a secret prompt” or “optimizer’s curse.”
I downloaded the current 627-entry VibeMathed dataset and recomputed totals by named human collaborator. Alpöge has seven entries whose significance scores sum to 383. Jihao Liu has fourteen summing to 286, Boris Alexeev thirteen summing to 162, Terence Tao nine summing to 150. So your observation is real in one important sense: Alpöge is decisively ahead in significance-weighted named collaborations, although he is not ahead in volume or number of results classified as AI-discovered.
But “significance” is an editorial, AI-generated estimate of how famous or important the original problem was, not a measurement of the magnitude or reliability of the new result. VibeMathed explicitly says scores are honest only to about a five-point band, and that status, AI role and verification are separate axes. It also scores the problem before its solution, so a partial zeta result can outrank a claimed complete solution of the Hopf problem. This makes 68 versus 65 effectively a tie, not a meaningful three-point victory. See the site’s unusually careful methodology.
Here is the quality-adjusted corpus as of August 25:
Result | Publicly documented division of labor | Current evidential status | What it says about Alpöge |
|---|---|---|---|
Zeta zeros, score 68 | Jarred Sumner gave an unreleased Claude the target; Claude ran the search and found the argument. Alpöge and Ralph Furman entered afterward to understand, contextualize and check it. | Genuine partial advance, extensively checked and Lean-formalized. | Strong evidence that Alpöge is an excellent validator and mathematical interpreter, almost no evidence about his elicitation skill. Anthropic’s full process account is explicit about this. |
Complex structure on (S^6), score 65 | Alpöge publicly identifies the triangle-group and torus-family construction as something he loves; Opus 5 produced the 100-plus-page detailed manuscript. The discovery trace is undisclosed. | Two-day-old, unreviewed candidate that directly conflicts with a published corrigendum. No formalization or independently checkable finite certificate. | Potentially the strongest evidence of a real human-model synthesis, if it survives. The conceptual construction overlaps strikingly with Alpöge’s prior interests. See the 108-page paper and current MathOverflow discussion. |
Jacobian conjecture, score 65 | Akhil Mathew suggested looking for counterexamples; Fable found the explicit map after Alpöge prompted it. Anthropic says a persistence or encouragement prompt similar to the zeta run was involved, but the trajectory and failed attempts remain private. | The counterexample is real and trivial to check once supplied. It disproves the conjecture in dimensions three and higher; dimension two remains open. | The cleanest major win. It proves that Alpöge can place a public frontier model on an exceptionally valuable, tractable target and close the loop in hours. Tao’s mathematical digestion and Gao’s subsequent generalization make clear that this was not a press-release mirage. |
Smooth Carathéodory conjecture, score 55 | Achim Wanner supplied the problem socially. Alpöge says he saw the complicated literature, asked whether it was acceptable to try a model, and then he, John-Paul Smith and Claude checked the construction. The wording most naturally credits the construction to Alpöge rather than Claude. | Candidate, with a compiling Lean artifact whose correspondence to the original statement has not yet received an independent audit. | Strong evidence for problem brokerage, rapid cross-field learning and responsible pre-circulation. Weaker evidence for AI discovery. See Alpöge’s announcement and the linked artifacts in the VibeMathed dataset. |
Elliptic-curve rank 30 and 31, scores 50 each | Found by Claude, Alpöge and Ava Howell, with no published process account. | The rank-31 leaderboard displays 31 independent points; the stronger exact-rank statement is conditional on BSD and GRH. This is one rapidly advancing search program, not two independent successes. | Strong domain fit and computational-search evidence, but public provenance is thin. See the ICARM rank-30 announcement and rank-31 record. |
Hadamard order 668, score 30 | Alpöge, Philippe Voinov, Saul Reynolds-Haertle and Claude produced constructions filling all previously unknown admissible orders below 2000. Alpöge jokingly claimed only “bad suggestions”; the actual division of labor is undisclosed. | The matrices are exact, cheap to verify and accepted as solving the finite construction problem. Epoch provisionally classifies it as AI-solved pending fuller disclosure. | A textbook verifier-first target and evidence of good team formation, not clean evidence of Alpöge’s personal discovery rate. See Epoch’s record and Ion Nechita’s useful contemporary assessment. |
So seven nominal hits reduce to something like three or four statistically distinct Alpöge-operated programs:
Jacobian.
The elliptic-rank search.
Hadamard, with unclear credit.
The (S^6) and Carathéodory candidates, if they survive.
The zeta discovery should not be counted as an Alpöge elicitation hit, and rank 30 plus rank 31 should not be treated as independent draws. Conversely, merely reducing the list this way should not obscure how extraordinary the residual is. One definitive landmark counterexample, a record-search program, a finite construction result, and two serious candidates within roughly five weeks is still an enormous outlier.
If the (S^6) argument survives, I agree with your commenters that it will probably feel more consequential than the zeta partial result. VibeMathed’s numerical ordering is not designed to capture that distinction.
I suspect the salient skill is not “prompt engineering” in the popular sense. The best evidence suggests that prompts are becoming surprisingly generic. The scarce human ability is constructing the whole problem-to-publication pipeline.
Most of his results have the form “hard to find, cheap to falsify or verify”:
An explicit noninjective polynomial map.
A finite matrix satisfying an exact identity.
Independent rational points on an explicit curve.
A support function with computable differential properties.
A complex-geometric construction whose decisive topology can, at least in principle, be reduced to finite matrix and homology calculations.
These are unusually suitable for a fallible but enormously persistent generator. The model can search an absurd hypothesis space, while algebra systems, Lean, exact arithmetic or a human expert can kill false candidates cheaply.
This is more subtle than simply choosing famous open problems. Asking Claude to prove arbitrary famous theorems would produce a torrent of elegant nonsense. Alpöge appears to identify famous problems whose obstruction might be representational or search-based, rather than requiring an entirely unavailable theory. That is mathematical “interface design”: finding a representation where model generation and exact verification meet.
His personal site and CV show serious work across analytic number theory, arithmetic statistics, elliptic curves, combinatorics, finite groups, effective Diophantine geometry and computation. He was using CUDA for medical computer vision in high school, then became an elite pure mathematician, then describes GPT-4 as the thing that pulled him back toward computer science.
There is also a revealing match between the new results and his older expertise. His CV records a 2022 talk on triangle groups, Belyi uniformization and modularity. The (S^6) candidate is built around the ((3,4,\infty)) triangle group and a universal family of tori. That is not evidence of a random employee feeding random famous problems to Claude. It looks like a model elaborating a construction at an interface where Alpöge already had strong taste.
The out-of-domain problems arrive through a different channel. Akhil Mathew suggested the Jacobian target. Achim Wanner supplied Carathéodory in a conversation about favorite problems. Ava Howell, Voinov and Reynolds-Haertle join specialized searches. So the operating model is:
flowchart TD
A[“Breadth and credibility”] --> B[“Experts send valuable problems”]
B --> C[“Verifier-aware model portfolios”]
C --> D[“Fast checkable results”]
D --> E[“More trust and collaborators”]
E --> B
That is an increasing-returns flywheel. After the Jacobian result, Alpöge became a Schelling point for “I have an old problem that might now be vulnerable.” A random equally skilled prompter does not receive that inbound target flow.
The Morgan Prize is awarded annually for outstanding undergraduate mathematical research in the United States, Canada and Mexico. It is not literally an exam ranking that identifies “the best undergraduate mathematician in America.” Alpöge’s broader record is even more diagnostic: 4.0 at Harvard, highest GPA, best senior thesis, a physics master’s alongside mathematics, six undergraduate research papers, Princeton under Manjul Bhargava and a Harvard Society of Fellows appointment. The official prize description and his CV support the underlying claim that he was an extreme mathematical outlier unusually early.
Why does this matter for centaur performance? Current models generate huge quantities of locally plausible material. An elite mathematician can:
Recognize when an approach is genuinely new rather than terminological remixing.
Detect the one promising branch among hundreds.
Supply a useful normalization or representation.
Know which objection will immediately occur to an expert.
Compress a giant model artifact to a two-page mathematical kernel.
Put their reputation behind the result when it is ready.
The bottleneck is moving from generation to justified belief. Morgan-level talent is extremely useful there even when the model supplied the decisive object.
The public chronology is revealing: the Jacobian run during a World Cup final, Hadamard as “weekend fun,” five Astra reproductions within 24 hours, two elliptic-rank records in three days, then an (S^6) manuscript immediately afterward.
This suggests an enormous attempt denominator. He seems to regard frontier models not as an occasional assistant but as a continuously running research substrate. Most mathematicians with Fable access have neither the habit, compute allocation, target backlog nor willingness to launch dozens of unreasonable attacks.
That intensity can masquerade as an inexplicably high hit rate when we observe only successful announcements. Alpöge may have a moderately better conditional hit rate and a vastly larger number of well-chosen serious attempts.
He announces concise, checkable objects rapidly, often before a conventional manuscript exists. That has costs, as the Hadamard criticism and present (S^6) skepticism show. But it also gets experts checking the result within hours and ensures it enters trackers such as VibeMathed.
A slower mathematician might spend three months preparing a traditional paper and be invisible during the window in your screenshot. The leaderboard therefore measures something like discovery, validation, communication and willingness to claim public priority. Alpöge is unusually optimized across all four.
There is also an attribution-policy confound. OpenAI’s Astra announcement says the mathematical arguments were generated by Astra, humans prepared the manuscripts, and it would be misleading to claim human authorship. It names no individual “centaur” for those ten results. Anthropic’s zeta paper instead names Alpöge and Furman as the responsible mathematicians even though Claude and Jarred Sumner found the result. Thus a leaderboard based on named collaborators mechanically credits Alpöge while rendering some competing internal operators anonymous. OpenAI’s policy is stated in its ten-advances announcement.
A lot, but probably less than the strongest access-only story implies.
Anthropic confirms that its unreleased internal Model 2 is used extensively for research and engineering, interactively and in persistent agent deployments. It characterizes Model 2 as noticeably better than Mythos 5 for some internal tasks but only slightly better overall, and not a jump comparable to the earlier Opus-to-Mythos transition. See the August risk report.
Because Alpöge works at Anthropic, it is reasonable to infer that he may have access. But there is no public evidence that Model 2 produced any result in his current streak:
Jacobian is credited to public Fable.
(S^6) explicitly mentions Opus 5 for the long write-up.
Zeta used an unreleased research Claude, but Jarred Sumner operated the discovery run.
Hadamard and the rank searches disclose no version.
Carathéodory names Claude but not a version.
So “Model 2 caused the streak” is presently speculation.
The more revealing quasi-experiment is Alpöge’s response to OpenAI’s ten Astra results. Once OpenAI disclosed which targets Astra had solved, Alpöge reported that publicly available Fable reproduced problems 4 through 8 within 24 hours, autonomously, with generic prompts and no internet. His original post and Gary Marcus’s analysis capture the point.
That observation simultaneously implies three things:
First, exclusive model quality was not sufficient to explain Astra’s apparent lead.
Second, knowing which problems are currently soluble is tremendously valuable. OpenAI had paid the exploration cost across an undisclosed set of successes and failures; Alpöge was handed five positively selected targets.
Third, Alpöge still deserves operator credit for being the person who immediately recognized and exploited this experiment. Many people had public Fable. Almost nobody responded by launching a paranoid, leakage-controlled replication campaign overnight.
“Everyone has the model” is also operationally false. They may have the same public weights, but not the same:
Rate limits and willingness to spend.
Persistent-agent infrastructure.
Internal model experience before release.
Colleagues who know model failure modes.
Literature and exact-computation tooling.
Ability to run dozens of branches while treating failures as disposable.
Expert network for immediate validation.
The marginal inference cost is already surprisingly affordable. Fable costs $10 per million input and $50 per million output tokens, according to Anthropic. The zeta run used 31 million output tokens, roughly $1,550 in output charges at those public rates, before input and caching. OpenAI says its ten successful Astra discoveries would total about $2,000 at Sol API rates. The hidden quantities are failed-run compute, infrastructure and expert time. This is not an insurmountable university-scale capital advantage, but it is a substantial advantage over ordinary subscription chat usage.
Suppose each person has a latent success probability (p_i), runs (T_i) attempts, and we observe only public successes. Selecting the largest observed result from hundreds of users means the winner’s apparent (p_i) will be inflated by noise. We should expect regression toward the mean and probably a different leaderboard topper eventually.
But the Alpöge observations violate the simplest lucky-winner model in several ways:
The hits are mechanistically related to his background and network.
He demonstrates abnormally high effort and response speed.
He repeatedly selects verifier-friendly targets.
He has access to unusually strong infrastructure.
The wins create target-flow and validation advantages that make future wins more likely.
The more serious statistical issue is not random luck; it is denominator blindness and dependence. We do not know how many failures he ran. The seven entries are not seven independent Bernoulli successes. They share Claude infrastructure, collaborators, problem lists and one compressed period of model capability. Zeta is not his search, the two ranks are one campaign, and two of the highest-ranked claims remain candidates.
My subjective calibration would be:
Probability that he is currently the leading publicly visible, significance-oriented human AI-math operator: above 90 percent.
Probability that he has the highest intrinsic per-attempt success rate among all serious operators, including unpublished lab personnel: perhaps 50 to 65 percent.
Probability that ordinary winner’s luck is the dominant explanation of the observed gap: below 20 percent.
Probability that the exact same person is still a clear public number one in twelve months: below one-third.
The last figure is not skepticism about Alpöge. It is a forecast about a violently expanding competitor pool.
The gameboard flip changes the objective function.
The previous prestige game heavily rewarded spending years proving one deep theorem in a narrow specialty, then communicating it through established institutions. The emerging game rewards:
Maintaining a portfolio of dozens of potentially vulnerable questions.
Knowing which questions can be converted into constructions, searches or finite certificates.
Allocating enormous cheap cognitive labor without attachment to failed branches.
Rapidly recruiting domain experts and formalizers.
Separating the two-page mathematical idea from the 100-page verification artifact.
Publishing at the speed of verification rather than the speed of journals.
Alpöge is unusually well-shaped for both regimes. He has enough old-regime capital to judge results and command attention, but unusually little attachment to old-regime workflow. His site even says that ceasing to publish solo papers in journals after undergraduate years was deliberate. That combination is rarer than either elite mathematics or model access alone.
This is analogous to early computing revolutions in experimental science. The first winners are not necessarily the very best theoreticians or the very best programmers. They are people with enough theory to select consequential questions, enough computational instinct to reformulate them, and enough social capital to turn machine output into accepted knowledge.
These are explicitly rough, falsifiable forecasts.
Horizon | Operating model | Scarce advantage | Alpöge forecast |
|---|---|---|---|
6 months, February 2027 | A top operator maintains a private problem registry and runs many asynchronous projects. Each target has a verifier, kill criteria, a literature bundle and several adversarial reviewer agents. Humans spend more time triaging dashboards than composing prompts. Explicit counterexamples, finite constructions, records and bound improvements remain overrepresented. | Problem selection, compute allocation, exact verification and access to experts who can audit statement fidelity. | Roughly 65 to 75 percent that he remains top five among named individuals; around 40 percent that he is still the unambiguous number one. Whether (S^6) survives will materially move both estimates. |
12 months, August 2027 | Stateful mathematical workbenches become normal among top groups. Literature search, CAS experiments, Lean formalization, branch history and independent reviewer loops are integrated. A “centaur” increasingly means a small research cell, not one person chatting with one model. | Proprietary problem backlogs, trusted expert networks, unpublished context, and the ability to digest thousands of correct but unimportant outputs. | Likely still a prominent operator, but the public leaderboard becomes much more volatile. I would put his probability of literal number one around 20 to 30 percent, while expecting him to remain top ten. |
18 months, February 2028 | Models handle most branch generation, computational experimentation, write-up and formal verification. Humans commission research programs, decide what is worth knowing, audit the problem statement and explain why a result matters. In the faster capability scenario, “centaur mathematician” already sounds too individualistic. | Agenda-setting, semantic judgment, institutional credibility and ownership of the verification pipeline. | Around 60 percent that an individual leaderboard is no longer the natural unit. Alpöge’s durable advantage would be as the nucleus of a high-output research network, not necessarily as the named collaborator on the most entries. |
Google’s recently described AI co-mathematician is already a prototype of this workflow: an asynchronous, stateful workspace with parallel branches, persistent failed hypotheses, reviewer loops and standard commercial models. Its orchestration increased FrontierMath Tier 4 performance from 19 percent for the base Gemini model to 48 percent, although with unrestricted inference and higher cost. The same paper warns that model reviewers can converge on a flawed consensus that neither the agents nor humans easily detect. See the AI co-mathematician paper.
That tells us where the competitive frontier is moving. Prompt phrasing becomes commoditized. Research orchestration, uncertainty management and statement auditing become the differentiators.
The crucial evidence is not another impressive tweet. It is process disclosure and precommitment.
Evidence that would strengthen the “Alpöge has discovered a genuinely general centaur method” thesis:
The (S^6) paper survives sustained review with its central construction intact.
Full Jacobian or future transcripts show a generic target plus limited human mathematical steering.
He produces several more high-significance results from a predeclared problem set where failures remain visible.
Outside mathematicians using the same public model and targets fail to reproduce his success despite similar compute.
His results keep spanning new fields without depending on a friend first supplying exactly the vulnerable problem.
Evidence that would narrow the thesis:
(S^6), Carathéodory or the rank claims require major correction.
The private transcripts reveal highly developed human ansätze or extensive undisclosed assistance.
Public model workbenches make his apparent hit rate commonplace within months.
Attribution-normalized data show several anonymous lab teams already matching his output.
His future wins remain concentrated in publicly revealed targets, making him an exceptionally fast follower rather than a uniquely good target discoverer.
The single most informative missing statistic is an attempt ledger: problems tried, compute spent, how targets were sourced, when humans intervened, and why runs were stopped. Without that denominator, “hit rate” is a metaphor rather than a measurable quantity.
My central model, then, is not “Levent knows magic words for Claude.” It is that he may be the first elite mathematician to operate as a mathematical venture capitalist, systems engineer and research closer at once. He gathers unusually valuable deal flow, recognizes what the new machinery can attack, runs a large portfolio, kills failures cheaply, recruits verification and publishes decisive artifacts at very low latency. Access and timing put him in position; exceptional mathematical judgment and an emerging reputation flywheel turn that position into a visible lead.
That looks like a genuine paradigm-specific winner. It also contains the mechanism by which a new winner could overtake him quickly: once the workbench is commoditized, the contest shifts again toward whoever owns the best problem flow, verification organization and research agenda.
A monthly re-audit would be useful here because the (S^6) review and leaderboard could change the conclusion quickly.
Was differentially funding this a focus of your microgrants last time?
I’ve been told explicitly, for example, that “for this project to go bigger you should avoid the LessWrong brand like the plague.” I’m currently confused if this is an incorrect personal bias from that person, a narrowly correct statement about messaging to academics, or a broadly correct statement about messaging to the general public.
This tangentially reminded me of Ryan’s observation, albeit about DC policy folks
LessWrong is famous in these policy circles. Even out here in DC. “Of course, we’ve all read LessWrong,” started one speaker. I sensed a kind of grudging respect for those “abrasive,” “insular,” and “truth-telling” rationalists of SF (okay, the “truth-telling” quote is from me but I got the sense that many people value what you all are putting out on here).
I might write up how Williamson’s metaphilosophical anti-exceptionalism implies we should automate philosophy.
I’d be keen to read this if you ever got around to writing it up.
I was confused why this was ever a thing. I just assumed everyone had seen this chart and noticed how even if the blue line plateaus the green and red need not, especially given tremendous persistent economic incentives. Maybe the counterargument is “obviously the proposers knew this, what they actually proposed were thresholds that would ratchet downward over time to adjust”?
(“the implicitly invoked norm that every comment on a post needs to respond super directly to the content of a post” seems a bit strawmannish? The “it being the most-upvoted and first-appearing comment” conditional seems key to his frustration, I’m guessing he wouldn’t be as annoyed if it weren’t first-appearing and merely highly upvoted.)
The report said they “do not currently have plans to release this model externally”, so not anytime soon if at all I suppose. It also concluded (like you said) that there doesn’t seem to be any trend acceleration yet.
I just learned about Anthropic’s internal Model 2.
… somewhat more capable than Mythos 5. Our rough qualitative sense is that this model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview. We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities.
… based on limited data, Model 2 (not shown on this figure) appears to be around 1.5 points higher on AECI than Claude Mythos 5, with large error bars (a smaller increase than the one from Claude Mythos Preview to Claude Mythos 5)
(brought here by Kaj Sotala’s comment) I’d be curious how you’ve updated on these takes. My own vague impressions: still bad at research taste, reliability has improved tremendously but still brittle in some sense, lots of progress solving non-templated problems especially in math but even there original seeing seems missing, still considerably diminishing returns on problem-solving time. Overall not really closing the gap to Byrnesian AGI, although maybe closing the gap to PASTA-AGI.
Can you say more on those theories, or share pointers?
On Dec 13th 2017, a still-anonymous redditor announced they would donate 5,057 BTC, the majority of their wealth, to charitable causes via the Pineapple Fund.
I remember staring at bitcoin a few years ago. When bitcoin broke single digits for the first time, I thought that was a triumphant moment for bitcoin. I watched and admired the price jump to $15.. $20.. $30.. wow!
Today, I see $17,539 per BTC. I still don’t believe reality sometimes. Bitcoin has changed my life, and I have far more money than I can ever spend. My aims, goals, and motivations in life have nothing to do with having XX million or being the mega rich. So I’m doing something else: donating the majority of my bitcoins to charitable causes. I’m calling it 🍍 The Pineapple Fund.
Yes, donating ~$86 million worth of bitcoins to charities :)
The Pineapple Fund (“just Pine and a friend”) ended up receiving over 10,000 applications and donating 5,104 BTC ($55.75M) over 4 months to 60 charities: $5M to GiveDirectly, $5M to the Open Medicine Foundation, $5M to MAPS for psychedelics research, $2M to charity: water, $2M to the SENS Foundation, etc.
I asked Sol to estimate how much impact 🍍 had, by doing “lookbacks” with a pluralistic accounting of “impact”, since Pine’s giving was in the spirit of Richard Ngo’s “things that I personally am unusually excited about”. Its one-sentence summary was
My median view is therefore: the fund was substantially but not fully additional, unusually good at creating runway and durable capacity in small organizations, produced several large concrete service outputs, and created genuine research/public-infrastructure option value.
Sol’s analysis (all 60 grants individually & overview), based on work in this spreadsheet
Evidence cutoff: 16 August 2026
Portfolio: 60 grants made from December 2017 to March 2018
Purpose: a retrospective evaluation, not a recommendation about where to donate now
There is no honest single number for the Pineapple Fund’s impact. The strongest answer is a vector of outcomes plus a separate attribution model:
Ledger: Pine’s 60 displayed grant labels sum to 55.75m. The often-cited 55m is rounded. A third-party ledger says 55.6m only because it records Reagent (100k) and Enthea ($50k) as $0; I do not silently discard those two grants. The fund reported 5,104 BTC ultimately donated, versus 5,057 BTC in the launch announcement.
Counterfactual mission resources (not social impact): my grant-by-grant model says roughly 26, 485, 000–49,536,500, central $39,219,500, probably increased grantees’ mission resources relative to the no-Pine world. That is 48%–89% of nominal dollars, central 70%. This is a deliberately broad sensitivity range based on gift restriction, grantee size, evidence of a binding funding constraint, absorption/delay, and other funders.
Best same-unit welfare result: GiveDirectly’s $5m equals 15,823 simple $316 annual-doubling equivalents. Applying GiveWell’s updated model mechanically gives 28,680 recipient consumption-doubling units plus 18,150 nonrecipient spillover units = 46,830 consumption units, before my 80%-100% grant-additionality discount. I keep GiveWell’s modeled mortality and ‘other’ benefits separate because adding them would import moral weights.
Best same-unit mortality benchmark: New Incentives’ $250k would equal about 42–167 deaths averted, central 118, at mature-program costs of 1, 500−6,000/death (GiveWell’s retrospective point estimate is $2,117). But this is not a clean 2018 lookback: GiveWell had already committed $5.944m to fund the RCT period. After a 20%-80% grant-specific additionality discount, my highly uncertain Pine-attributable benchmark is roughly 8–134 deaths, central 59—about 0.4k–9.4k DALYs at 50-70 DALYs/death.
Water: The Water Project directly identifies 37 projects. A project-size model gives about 15k–25k people initially served. charity: water plausibly adds 25k–45k projected people through the $1m field-program share. After separate attribution discounts, the combined estimate is roughly 25k–68k initial/projected water users, central about 43k. Long-run reliable-water person-years are much less certain.
Health services: Watsi plausibly represents 2k–5k treatment budget-equivalents (about 0.9k–4.5k after attribution). Possible represents roughly 40k–70k service/catchment equivalents (about 12k–56k after attribution), not DALYs.
Housing, food, trees, devices: New Story plausibly financed 250–400 home-equivalents before attribution; Replate 300k–700k meal-equivalents; TreeSisters 0.5m–1.5m trees funded; human-I-T 1k–3k device/household equivalents. These remain output units, not welfare conversions.
Research: MAPS, OMF, SENS, Organ Preservation Alliance, Methuselah and related grants clearly bought trials, projects, centers and field-building. Yet realized population treatment impact through the cutoff is zero or not measurable. MAPS is the sharpest negative/deferral update: positive Phase 3 results and at least $4m of matching gifts, but FDA did not approve the therapy in 2024; a 2026 resubmission preserves option value.
Open infrastructure: Apache’s ’1m′becameabout * *892,882** and roughly 6–8 gross operating months; FSF’s became about $860k and 6–7 months; Internet Archive received about 41–52 operating-day equivalents plus a $1m match; OSM funded 12 microgrants after a two-year delay and had deployed only about one-third of the restricted grant by the end of 2021. These are more defensible than multiplying global user counts by Pine’s budget share.
Identifiable crowd-in: MAPS (at least 4m), InternetArchive(1m), The Water Project (1m), Watsi(500k) and Erowid (250k) togetherreportatleast * *6.75m** of follow-on or matched gifts. I assign only 25%-75% of that as counterfactually crowded in (1.7m−5.1m) and do not add it to the mission-resource total because doing so risks double counting and donor displacement.
My median view is therefore: the fund was substantially but not fully additional, unusually good at creating runway and durable capacity in small organizations, produced several large concrete service outputs, and created genuine research/public-infrastructure option value. It was also much less auditable than its on-chain transparency suggested. Blockchain transparency proves transfers; it does not prove liquidation value, deployment, counterfactuality, or outcomes.
Layer | Question | What counts | What does not count |
|---|---|---|---|
Transfer | Did Pine send it? | Original ledger row and transaction/recipient confirmation | The December 2017 $86m mark-to-market promise |
Realization | What cash did the grantee receive? | Disclosed conversion proceeds; otherwise headline nominal amount with a caveat | Assuming every $1m headline became exactly $1m cash |
Lookback | What was actually done? | Restricted projects, capital assets, trials, publications, audited spending and contemporaneous outputs | All later organization-wide reach |
Counterfactual resources | How much extra mission capacity did Pine create? | Explicit 3-point additionality range based on restriction, size, absorption and alternative funding | A claim that these dollars all caused outcomes |
Outcome model | What uplift is plausible anyway? | A disclosed denominator, a gross output pool and a second attribution discount | Cross-domain summation or hidden moral weights |
Downstream | Did the grant unlock later activity? | Named matches, successor institutions and clearly timed scale—reported separately | Counting every descendant outcome as Pine’s |
The user specifically asked me not to let ‘unresolved’ collapse to zero. I therefore do both analyses for each grant. The lookback records what can be traced. The counterfactual model then makes the best outside-view estimate I can justify, including for fungible general support. They are shown in different workbook sheets and different paragraphs below.
The initial Reddit post valued 5,057 BTC at about $86m using a spot price of $17,539/BTC. That was a promise valued at a market peak, not cash received by charities. Pine’s farewell account says 5,104 BTC produced roughly $55m after volatility. The original website’s 60 USD labels sum to $55.75m. For impact calculations I use the row labels as the common portfolio ledger, then substitute disclosed realization only in grant-specific calculations. This avoids pretending to know 60 sale prices.
Most interviewed recipients said they sold quickly, but there were operational and KYC delays; at least one large recipient held for about two weeks. The consequences are visible: Apache’s 88.34 BTC was later recognized at $892,882, and FSF says its 91.45 BTC became about $860k by conversion. MAPS moved the other way: it reports about $5.339m from Pine after a volatility/fee top-up. The correct uncertainty is therefore recipient-specific, not a blanket haircut.
Twenty-nine grants of at least 1maccountfor * *49,250,000 (88%)**. Grants of $2m or more account for $31m (55.6%). The small-dollar tail is numerous but financially minor; however, it contains several of the least transparent grants and therefore matters disproportionately for process learning.
Broad bucket | Nominal amount | Share |
|---|---|---|
Medical & scientific research | $16,650,000 | 30% |
Poverty, livelihoods & inclusion | $9,700,000 | 17% |
Rights, justice & equity | $7,250,000 | 13% |
Digital & public infrastructure | $7,050,000 | 13% |
Education | $4,000,000 | 7% |
Global health delivery | $3,250,000 | 6% |
WASH | $3,000,000 | 5% |
Environment & conservation | $2,400,000 | 4% |
Animal welfare | $1,250,000 | 2% |
Health support & advocacy | $1,200,000 | 2% |
Evidence grade | Grants | Nominal amount | Share |
|---|---|---|---|
A | 10 | $15,000,000 | 27% |
B | 24 | $22,050,000 | 40% |
C | 16 | $12,150,000 | 22% |
D | 9 | $6,500,000 | 12% |
U | 1 | $50,000 | 0% |
Grades A+B cover $37,050,000 (66%). That does not mean two-thirds of dollars have proven outcomes: B often means a convincing account of capacity or use, not a causal health/income effect.
GiveDirectly. The simple calculation the user proposed is:
$5,000,000 ÷ $316 = 15,823 annual-doubling equivalents.
GiveWell’s October 2025 Rwanda illustration for a $1m donation instead yields 5,736 recipient consumption units, 3,630 spillover consumption units, 837 mortality ‘units of value,’ and 818 other units. Keeping only the consumption-denominated components:
5 × (5,736 + 3,630) = 46,830 consumption-doubling welfare units.
That central estimate is about 3.0 times the simple 15,823 figure, which is consistent with GiveWell’s broader 3-4x revaluation. Two live upside channels remain: using Egger et al.’s spillover estimate at face value would put spillovers near 180% of recipient benefits rather than GiveWell’s adjusted 60%-70%; and preliminary 5-7 year follow-up suggests more persistent consumption gains. Either moves the program toward 4-6x the old estimate. Important downside channels remain too: external validity, imprecise spillovers, distribution of gains toward richer local households, country/program mix, and the unknown split between lump-sum transfers and the 12-year UBI experiment.
I apply an 80%-100% grant-additionality range, yielding roughly 28k–65k Pine-attributable consumption units, central about 43k. This extra discount is about Pine’s grant, not GiveWell’s causal effect estimate.
New Incentives. GiveWell’s corrected 2025 lookback on a completed $16.8m 2021-23 grant estimates 7,937 deaths averted, or $2,117/death, versus an initial $3,910/death. Costs fell from about $38 to $18 per enrolled child; GiveWell also raised estimated vaccine-preventable mortality. A mechanical application to Pine is:
$250,000 ÷ $2,117 = 118 deaths averted.
My 1, 500−6,000 sensitivity gives 42-167 deaths. But the critical lookback fact is that Pine’s February 2018 money arrived after a $5.944m GiveWell incubation grant intended to fully fund the RCT program through May 2020. Pine may have bought resilience, complementary work, or runway into scale rather than an extra $250k of mature vaccinations. I therefore apply a wide 20%-80% grant-specific additionality range, producing 8-134 deaths, central 59. At 50-70 DALYs per death that is roughly 0.4k-9.4k DALYs, central about 3.5k. This is the most useful speculative estimate—and must not be mistaken for a grant-level observed result.
MAPS and other biomedical research. I book zero realized population treatment DALYs through the cutoff. That is not a claim of zero value; it is a refusal to convert research option value into DALYs without approval, adoption, durability, and counterfactual treatment evidence. Phase 3 trial arm differences imply perhaps 20-35 incremental short-term participants no longer meeting PTSD criteria across the trials; a financing-share/acceleration model gives roughly 2-18 Pine-attributable trial responses. Those are research outcomes, not public-health impact. OMF, SENS, OPA and Methuselah remain in the same option-value bucket.
Watsi: $21.5m/34,887 lifetime patients = $616.27 per patient; 2m/616.27 = 3,245 treatment budget-equivalents, plausible gross range 2k-5k and Pine-attributable range about 0.9k-4.5k.
Possible: contemporaneous unit costs were about $20.56 per capita and $14.76 per visit. Thus $1m corresponds to 48.6k catchment-person-years or 67.8k visits. I use a 40k-70k gross bracket and 12k-56k after attribution. These are services, not DALYs.
PTSD Veteran Athletes: a public ~$5k/participant benchmark gives about 10 participant-program equivalents, plausible Pine-attributable 4-13. Outcome evidence is testimonial.
TalkLife, Empower Work, Experience Camps and Canine Therapy Corps: I estimate user/contact/camp/hour uplift in their own units in the grant appendix; no controlled health conversion is available.
The Water Project is unusually concrete: 37 grant-tagged projects. At 400-675 people per project, that is 15k-25k initial users, central 18.5k. If they functioned for four to seven years, a rough 60k-175k reliable-water person-years would follow; I use 60k-150k to stay conservative. Its service hub, water-quality lab, spares, training, and monitoring plausibly improve durability but are exactly why a simple construction count can understate impact.
charity: water reported a 1mfield/1m operations split. In 2017 it deployed $34.1m for 1.184m projected people, or $28.8/person. Therefore 1m/28.8 = 34.7k projected users, bracket 25k-45k. Unlike The Water Project, I found no Pine-tagged completion/functionality list.
After grant-specific attribution, I estimate 25k-68k combined initial/projected users, central about 43k, without summing person-years that are not equally observed.
Pencils of Promise: 1.25m/9.893m × 96,339 = 12,170 student-year budget-equivalents, range 5k-25k, Pine-attributable about 1.8k-21k. A Guatemala assessment’s null result counsels against assuming enrollment equals learning.
Wiki Education: 500k/1.875m × 15,000 = 4,000 student-years, plus about 3.1m words at the same proportional share. The grant was unusually additional because accounts show a revenue shortfall and delayed hiring/projects.
Many Hopes: care and school accounts imply 58-71 child-years for $250k; modeled Pine-attributable 18-81 after a wider bracket.
Quill: the $1m explicitly financed staff, AI and product development. Later reach is enormous, but I allocate only 0.1m-0.9m incremental student-use episodes to Pine, not learning gains.
Mona: the organization grew sharply but suffered a 10 BTC compromise and did not tag outputs. The workbook provides a 3k-70k speculative additional student-support-year range after attribution; treat it as low-confidence.
These are intentionally not collapsed into a single welfare metric:
New Story: 250-400 home-equivalents before attribution; roughly 1k-2k resident places at 4-5 people/home.
Replate: current two-meals-per-dollar benchmark → 500k meal-equivalents, plausible Pine-attributable 150k-644k.
The Adventure Project: 2020 spending implies about $1,081 per job; $1m ≈925 livelihood-job equivalents, plausible Pine-attributable 150-1,200. Job duration and income effects are not standardized.
human-I-T: 1k-3k device/connected-household equivalents, Pine-attributable about 450-2,700.
TreeSisters: 0.5m-1.5m trees funded, Pine-attributable about 0.25m-1.38m. I do not convert seedlings to surviving trees or carbon.
Tennessee River Gorge Trust: Pine was in the financing stack for 409 newly protected acres linking a 3,157-acre corridor; estimated Pine-attributable 20-270 acres, not 3,157.
Nuzzles: one mortgage fully retired on a 16,000 sq ft/100-acre ranch—an observed durable asset. The downstream animal-capacity model is intentionally speculative.
Canine Therapy Corps: one building/down payment is observed; a rough before/after capacity model gives 110-540 annual therapy hours attributable to Pine.
Program-effect estimates can move more than grant execution. GiveDirectly did what it was supposed to do, but the estimated welfare generated rose 3-4x because spillovers, mortality and persistence changed. A 2018 operational audit alone would miss most of the update.
Costs can fall radically at scale. New Incentives later enrolled about twice the children expected for the same grant because cost per child fell from ~38to 18. Yet applying that mature cost retrospectively to Pine without checking the 2017 funding gap would create a new error.
Positive trials are not approval. MAPS’s case shows how a seemingly fundable final-mile regulatory path can still fail on trial design, safety, conduct and regulator confidence. The correct lookback is not ‘nothing happened’—Phase 3 and a resubmission happened—but public treatment impact is still zero.
Research capacity is easier to verify than research value. OMF expanded projects; SENS bought roughly a research-year; OPA helped create centers. Converting those inputs into DALYs today would mainly encode subjective probabilities.
Restricted money can still absorb slowly. OSM’s two-year delay and partial 2021 deployment show that earmarking is not the same as immediate additional activity.
Crypto created a hidden realization lottery. Apache and FSF lost 11%-14% versus headline; MAPS received a top-up. On-chain transfer transparency did not standardize cash value.
Small organizations often give the cleanest counterfactual narrative. HHR received more than one year of spending; Wiki Education had a documented shortfall; Nuzzles retired a specific mortgage; Canine bought a facility; Quill doubled down on product and staff. These do not prove outcomes, but they strongly narrow the no-Pine world.
Matches crowd in money, not necessarily impact. At least $6.75m of follow-on/matched gifts is identifiable, but some donors would have given anyway, some giving moved forward in time, and the funded activities can still fail.
Later reach is a poor default attribution rule. Let’s Encrypt’s billion certificates, Apache’s billions of users, OpenMRS’s 22m patients, and Quill’s 12m students are context. The right grant unit is often runway, capital, a project launch, or an incremental growth slice—not total reach times a budget percentage.
Missing reporting is itself informative. Indigo, Snow Lovers, Reagent, Neural Archives, Focus and parts of Alter are not scored zero. They receive broad modeled ranges, but the inability to resolve use nearly nine years later is a real downside relative to grants with explicit project ledgers.
For every grant I estimate a low/central/high counterfactual resource share. It answers: “What fraction of Pine’s nominal gift likely increased mission resources compared with the world where Pine did not give?” It is not a discount for program effectiveness; that comes later, if a unit conversion is possible.
The model uses five signals:
Restriction and traceability: direct transfers, named projects, capital purchases and matches get higher shares.
Gift size relative to organization: a grant exceeding annual expense gets a higher share than days of a nine-figure budget.
Binding constraint: documented shortfalls, delayed hiring, debt and explicit runway increase additionality.
Absorption and realization: BTC losses, delayed deployment, large reserves and opaque asset holding reduce it.
Alternative funding: pre-existing full funding, diversified corporate sponsorship or large simultaneous grants increase funging risk.
Summing grant × share gives $26,485,000 / $39,219,500 / $49,536,500. The central 70% figure should be read as an uncertainty-organizing device. Its value is that it forces assumptions for Apache, ACLU, Indigo and every other fungible grant into the open; it is not a frequentist confidence interval.
Where a same-unit outcome conversion is possible, I apply a second impact attribution factor to a gross output pool. This prevents a mistake such as treating Quill’s full post-2018 growth as Pine’s, or treating New Incentives’ mature 2021-23 cost per death as the literal use of a 2018 grant.
No recipient survey was conducted for this report. Public reports are uneven and often self-serving. A targeted data request to the largest 15 grantees could materially improve realization, restrictions and deployment.
Nominal dollars span volatile dates. I do not inflation-adjust because the relevant program-cost denominators are also drawn from different years; a half-adjusted portfolio would look more precise and be less correct.
Survival bias: organizations with good current websites are easier to trace; quiet projects may have failed or simply report poorly.
Attribution ranges are correlated. A world with high funging may also have lower program marginality. Do not treat row ranges as independent Monte Carlo draws.
Unit costs are averages, not margins. Watsi, charity: water, Pencils, Possible, Replate and others mix fixed and variable costs.
Outcome duration is mostly missing. Water functionality, housing occupancy, job persistence, treatment durability, tree survival and software security benefits are not harmonized.
Crowd-in can displace elsewhere. Matching gifts may reduce donations to other charities, an opportunity cost not modeled.
Research option value has fat tails. Booking zero realized DALYs understates expected future value if a therapy succeeds; assigning probability-weighted DALYs would require assumptions the user asked me not to hide.
Legal and public-infrastructure benefits resist unitization. I use operating capacity rather than pretending that user scale reveals causal impact.
Observed lookback — evidence B. The gift covered nearly a year of operations at a critical point; Watsi later reported 34,887 patients funded with $21.5m across 35 countries, and another anonymous crypto donor gave $500k after Pine. Strong capacity bridge; no patient ledger tying cases to Pine.
Separate counterfactual estimate. I assign 45%–90% additional mission-resource share (central 70%), or 900, 000–1,800,000 of incremental mission resources (central $1,400,000). Gift was unusually large and explicitly covered runway, but unrestricted funding can displace later fundraising and supported technology as well as treatments.
Gross modeled range: 2,000–5,000 treatment budget-equivalents (central 3,245). After the separate impact-attribution range, Pine-attributable uplift is 900–4,500 (central 2,271.5). Calculation: 2mdividedbyWatsi′sexactcurrentlifetimemean(21.5m/34,887 = $616.27 per patient) gives 3,245.3. Range allows case-mix, platform work, and displacement.
Downstream and forecast update. Plausibly preserved organizational momentum and helped normalize crypto giving; $500k identifiable follow-on is reported separately. Broadly positive, but the clean retrospective unit is weaker than the original ‘universal health care’ framing.
Main caveat. Lifetime mean cost is not 2018 marginal cost; later reach is not attributed to Pine.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence A. The organization identifies 37 Pineapple-funded projects in Kenya and Sierra Leone and says the money accelerated a service hub, laboratory, spare-parts system, training, and monitoring. A separate anonymous donor gave $1m within a week. One of the clearest grant-specific lookbacks; functionality and beneficiary counts remain incompletely published.
Separate counterfactual estimate. I assign 75%–100% additional mission-resource share (central 90%), or 750, 000–1,000,000 of incremental mission resources (central $900,000). Restricted, named projects and operational systems; downward adjustment for possible timing substitution and unverified long-run functionality.
Gross modeled range: 15,000–25,000 people initially served (central 18,500). After the separate impact-attribution range, Pine-attributable uplift is 11,250–25,000 (central 16,650). Calculation: 37 projects multiplied by 400-675 people/project (central 500). A separate durability scenario gives roughly 60k-150k person-years if systems function for 4-7 years.
Downstream and forecast update. The $1m follow-on gift is a credible catalytic effect, but only a fraction is counterfactually attributable. Delivery appears substantially realized; the main unresolved issue is durable service, not project construction.
Main caveat. People/project is modeled; do not equate initial access with continuously safe water.
Observed lookback — evidence B. The gift exceeded all funding BitGive had raised in its prior 4.5 years. GiveTrack and related work operated through 2022; BitGive reported work with nonprofits in 29 countries and more than 57,000 direct beneficiaries before transferring assets and IP to Heifer. Organizationally transformative and durable for several years; the stand-alone platform did not persist.
Separate counterfactual estimate. I assign 55%–95% additional mission-resource share (central 80%), or 550, 000–950,000 of incremental mission resources (central $800,000). Huge relative to prior funding and plausibly decisive for scaling, discounted for later donors, projects, and eventual wind-down.
Gross modeled range: 15,000–57,000 beneficiary episodes (central 30,000). After the separate impact-attribution range, Pine-attributable uplift is 5,250–48,450 (central 18,000). Calculation: Use BitGive’s 57k lifetime reach as an upper bound; assign 15k-57k to the post-grant platform era, then 35%-85% Pine attribution because the gift dominated prior capitalization.
Downstream and forecast update. Heifer received the technology and IP, so some option value survived closure; current donation functionality is not demonstrated. A real multi-year institution-building success, but less durable than a permanent crypto-philanthropy platform forecast.
Main caveat. 57k is organization-wide, self-reported reach, not independently verified welfare impact.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. EFF reported FY2018 expenses of $13.322m and a $3.946m surplus. The Pine grant equaled 7.5% of annual expense, about 27 operating days; no case or campaign was tied to the gift. Financial contribution is observable; mission outcome is not grant-attributable.
Separate counterfactual estimate. I assign 10%–60% additional mission-resource share (central 30%), or 100, 000–600,000 of incremental mission resources (central $300,000). Large, diversified organization with a strong surplus implies substantial funging, but unexpected unrestricted cash still expands reserves and risk tolerance.
Gross modeled range: 27–27 operating-day equivalents (central 27). After the separate impact-attribution range, Pine-attributable uplift is 2.7–16.2 (central 8.1). Calculation: $1m / $13.322m × 365 = 27.4 days; multiply by estimated counterfactual resource share.
Downstream and forecast update. Potentially supported litigation, policy, and technical work, but attributing EFF’s later wins would be false precision. Neither a clear disappointment nor a traceable success; primarily balance-sheet support.
Main caveat. Operating days are resource equivalents, not days EFF would otherwise have shut down.
Observed lookback — evidence A. The match was completed by more than 550 donors. Two Phase 3 trials produced positive symptom results; in MAPP2, 71.2% of MDMA participants versus 47.6% of placebo participants no longer met PTSD criteria. FDA issued a Complete Response Letter in 2024 and requested more work; approved-population treatment impact through the cutoff is zero. Research milestones achieved and donor match unlocked; the central promised endpoint—approval and access—has not yet occurred.
Separate counterfactual estimate. I assign 60%–100% additional mission-resource share (central 85%), or 3, 000, 000–5,000,000 of incremental mission resources (central $4,250,000). Explicit financing gap and matched Phase 3 work make funding highly additional, though other donors or delayed trials were plausible.
Gross modeled range: 20–35 incremental short-term trial participants no longer meeting PTSD criteria (central 28). After the separate impact-attribution range, Pine-attributable uplift is 2–17.5 (central 8.4). Calculation: Across roughly 190 randomized Phase 3 participants, trial-arm differences imply about 20-35 incremental short-term diagnostic responses. Pine attribution is limited to its financing share and counterfactual acceleration; this is not durable remission or public treatment.
Downstream and forecast update. At least $4m of matching gifts was mobilized. A 2026 NDA resubmission preserves option value, but no DALYs are booked before approval, uptake, and durable effectiveness. Large downward/deferral update versus 2018 expectations despite successful trials: regulatory, safety, blinding, and conduct risks dominated.
Main caveat. Trial response is not equivalent to cured PTSD; Pine-funded nonprofit work and later Lykos/commercial work overlap.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence D. The Foundation lists Pine at its 2017 Platinum contribution tier. No grant-specific developer, audit, release, or infrastructure output was found. Receipt is verified; deployment is unresolved.
Separate counterfactual estimate. I assign 35%–85% additional mission-resource share (central 60%), or 17, 500–42,500 of incremental mission resources (central $30,000). Small technical foundation and material sponsorship tier imply meaningful additionality, tempered by fungibility.
No defensible outcome denominator; the estimate stops at additional mission resources. Calculation: No output denominator exists; resource-share model only.
Downstream and forecast update. Likely supported hackathons, development, or infrastructure, but user/security benefits cannot be allocated. Unresolved rather than negative.
Main caveat. Do not multiply OpenBSD users by a speculative Pine share.
Observed lookback — evidence B. The grant was roughly half of SENS’s 2017 total expenses and close to one full year of its 2017 research spending. It expanded a research portfolio and runway; no approved rejuvenation therapy or measurable population health outcome resulted by the cutoff. Substantial research capacity purchased; realized clinical impact remains zero/unknown.
Separate counterfactual estimate. I assign 45%–90% additional mission-resource share (central 70%), or 900, 000–1,800,000 of incremental mission resources (central $1,400,000). Very large relative to research budget and flexible, but later crypto gifts and donors create substitution risk.
Gross modeled range: 9–13 2017 research-budget months (central 11). After the separate impact-attribution range, Pine-attributable uplift is 4.1–11.7 (central 7.7). Calculation: 2m/ 2.146m 2017 research expense × 12 ≈ 11.2 months; multiply by additionality range.
Downstream and forecast update. Research publications, spinouts, and successor capacity have option value; no responsible DALY estimate is possible without a causal chain to treatments. Downward on near-term health; ambiguous on scientific option value.
Main caveat. Research-budget months measure inputs, not probability-weighted longevity gains.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. The initial public update allocated $1m to field work and $1m to operations. No Pine-specific completion list was found. In 2017 charity: water reported $34.1m of project funding for 1.184m projected people served, a rough $28.8 per projected person. Program allocation is reported; completed beneficiary attribution is unresolved.
Separate counterfactual estimate. I assign 45%–90% additional mission-resource share (central 70%), or 900, 000–1,800,000 of incremental mission resources (central $1,400,000). Half was directly programmatic; the operating half has more displacement risk in a sizable organization.
Gross modeled range: 25,000–45,000 projected people served (central 34,700). After the separate impact-attribution range, Pine-attributable uplift is 13,750–42,750 (central 26,025). Calculation: $1m project share / $28.8 per projected person = 34.7k; range covers geography, monitoring, and cost variation. Operations are not converted again.
Downstream and forecast update. Reliable-water person-years could be much larger if projects remained functional, but no Pine-tagged functionality series was located. Likely delivered near expectations; evidence is weaker than The Water Project’s grant-specific reporting.
Main caveat. Projected people are not unique verified long-term users.
Observed lookback — evidence C. Mona reports growth from about 230k students supported in 2016 to 429k in 2019. GyanSetu expanded from 14 schools/490 students to 63⁄2,089 in 2020 and 100⁄2,698 in 2021, with 684 transitions reported. No Pine-tagged allocation is published. Strong post-grant growth and a disclosed crypto loss; causal allocation unresolved.
Separate counterfactual estimate. I assign 35%–85% additional mission-resource share (central 60%), or 350, 000–850,000 of incremental mission resources (central $600,000). Large gift to a mid-sized international education funder, but portfolio funding and later donors make displacement likely.
Gross modeled range: 15,000–100,000 additional student-support years (central 40,000). After the separate impact-attribution range, Pine-attributable uplift is 3,000–70,000 (central 16,000). Calculation: Use the 199k increase in annual student reach from 2016 to 2019 as an upper-bound growth pool; attribute 20%-70% of a conservative 15k-100k slice to Pine.
Downstream and forecast update. Expansion persisted well beyond 2018, but no learning-effect estimate is attached. Likely positive institutional uplift; less auditable than direct education grants.
Main caveat. Student reach is organization-reported and heterogeneous; crypto loss makes nominal-grant modeling optimistic.
Observed lookback — evidence C. Contemporaneous descriptions say the money was for homes for families. A mid-2010s New Story home was often cited near $6k; the organization later reported more than 20k people impacted, but no Pine-tagged home list was found. Mission-consistent delivery is likely; exact houses and occupancy are not traced.
Separate counterfactual estimate. I assign 55%–95% additional mission-resource share (central 80%), or 1, 100, 000–1,900,000 of incremental mission resources (central $1,600,000). Large early gift to a young organization and construction is concrete; allow for overhead, land, and other donors.
Gross modeled range: 250–400 home equivalents (central 333). After the separate impact-attribution range, Pine-attributable uplift is 137.5–380 (central 266.4). Calculation: $2m divided by 5k−8k per home. At roughly 4-5 people/home, this is about 1k-2k resident-place equivalents before attribution.
Downstream and forecast update. Institutional scaling may exceed construction count, but later fundraising and model changes dominate long-run reach. Probably near original practical expectations; evidence gap prevents a firm count.
Main caveat. Home cost is historical and does not include all enabling costs; ‘people impacted’ is not a housing outcome measure.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. Internet Archive publicly acknowledged both $1m gifts and met the second $1m public match. Against estimated annual spending around 14m−18m, Pine supplied roughly 41-52 gross operating days. No collection or preservation output was earmarked. Clear fundraising leverage and meaningful runway; content impact is not separable.
Separate counterfactual estimate. I assign 40%–85% additional mission-resource share (central 65%), or 800, 000–1,700,000 of incremental mission resources (central $1,300,000). Material gift with explicit match, but established donor base and unrestricted reserves create displacement.
Gross modeled range: 41–52 operating-day equivalents (central 46). After the separate impact-attribution range, Pine-attributable uplift is 16.4–44.2 (central 29.9). Calculation: $2m / annual expense × 365. Separately, apply 25%-75% counterfactual attribution to the $1m external match.
Downstream and forecast update. The Archive later rescued millions of Wikipedia links and continued preservation, but no share is assigned to Pine. Solid resilience and leverage rather than a measurable end-user outcome.
Main caveat. Operating days do not imply the Archive would otherwise have been offline.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. 2018 expenses were about $9.893m and reported reach was 96,339 students; Pine equaled 12.6% of expense. Outputs included 501 schools, 777 teachers, WASH work and e-readers. A Guatemala reading assessment found no statistically significant difference from comparison schools. Substantial program share; learning-effect evidence is mixed and grant allocation is not isolated.
Separate counterfactual estimate. I assign 35%–85% additional mission-resource share (central 60%), or 437, 500–1,062,500 of incremental mission resources (central $750,000). Material but fungible share of a diversified program; future donations could partially replace it.
Gross modeled range: 5,000–25,000 student-year budget-equivalents (central 12,170). After the separate impact-attribution range, Pine-attributable uplift is 1,750–21,250 (central 7,302). Calculation: $1.25m / $9.893m × 96,339 = 12,170. Wide range reflects program mix and fixed costs.
Downstream and forecast update. Education access and school quality likely improved for some students; no common learning-gain unit is defensible. Outputs delivered, but stronger skepticism is warranted about translating activity into learning.
Main caveat. Budget-equivalent reach is not a counterfactual enrollment or test-score gain.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. The grant reportedly financed expansion, cleanups, a mobile site, and credibility for other grants. Greensteps now reports roughly 8,000 bags plus 2,000 bulky items removed since 2017, but no Pine-specific count. Clear organizational inflection; environmental quantity is not grant-tagged.
Separate counterfactual estimate. I assign 55%–95% additional mission-resource share (central 80%), or 55, 000–95,000 of incremental mission resources (central $80,000). Early and reportedly transformative in a small organization; later volunteers and grants still matter.
Gross modeled range: 80–200 cleanup equivalents (central 140). After the separate impact-attribution range, Pine-attributable uplift is 28–170 (central 84). Calculation: Current sponsorship pricing around $500/cleanup yields a 200-cleanup ceiling; observed lifetime activity suggests using 80-200 before attribution.
Downstream and forecast update. Grant-eligibility and digital capacity may have crowded in later support. Small-grantee capacity bet appears positive, though effect size is imprecise.
Main caveat. Cleanup pricing is a current sponsorship benchmark, not 2018 marginal cost.
Observed lookback — evidence A. Cash was transferred through established programs, but the split and recipient ledger are not public. GiveWell’s updated 2025 model estimates cash transfers are 3-4x its prior estimate, with recipient consumption, spillovers, mortality and other benefits modeled separately. High-confidence causal program; grant-specific allocation is the main missing piece, not whether transfers occurred.
Separate counterfactual estimate. I assign 80%–100% additional mission-resource share (central 92%), or 4, 000, 000–5,000,000 of incremental mission resources (central $4,600,000). Direct transfers and a large experiment are highly spendable; GiveDirectly’s funding base creates some timing/funging risk.
Gross modeled range: 35,000–65,000 annual consumption-doubling welfare units (central 46,830). After the separate impact-attribution range, Pine-attributable uplift is 28,000–65,000 (central 43,083.6). Calculation: Simple historical comparator: 5m/316 = 15,823. Updated GiveWell Rwanda model: per $1m, 5,736 recipient + 3,630 nonrecipient consumption units; ×5 = 46,830. Range spans country/program mix and excludes mortality/other benefits.
Downstream and forecast update. GiveWell says full Egger spillovers or preliminary 5-7 year persistence would move cash to roughly 4-6x its old benchmark versus 3-4x now. Strong upward update: spillovers, mortality evidence, and possibly persistence were materially underestimated.
Main caveat. Current GiveWell marginal model is not a grant-specific 2017 allocation model; do not add its mortality ‘units of value’ to consumption units without moral weights.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. The two-week outdoor program is still active and has claimed roughly 100 veterans served annually; an external charity profile quotes about $5,000 per participant. No Pine cohort outcome data were found. Program delivery is plausible and survival is verified; health outcome evidence is testimonial.
Separate counterfactual estimate. I assign 55%–95% additional mission-resource share (central 80%), or 27, 500–47,500 of incremental mission resources (central $40,000). Small program and grant closely matched one cohort-scale budget; allow for other funding and fixed costs.
Gross modeled range: 8–14 participant-program equivalents (central 10). After the separate impact-attribution range, Pine-attributable uplift is 4.4–13.3 (central 8). Calculation: $50k divided by roughly 3.5k−6k per participant.
Downstream and forecast update. Potential PTSD, community, and employment benefits are not converted to DALYs because no controlled or standardized outcomes were found. Likely delivered a small cohort as intended; efficacy remains unknown.
Main caveat. Participant cost is a public benchmark, not audited grant accounting.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. Quill says the gift funded staff and technology that moved it to a new scale. It passed 1m students in 2018 and now reports 12m students, 8m in Title I schools, 42k schools and 3b sentences. A 2025 observational study found higher growth but only ESSA Tier III evidence. Explicit product/capacity use and strong scale; causal learning and Pine’s share of later adoption remain uncertain.
Separate counterfactual estimate. I assign 55%–95% additional mission-resource share (central 80%), or 550, 000–950,000 of incremental mission resources (central $800,000). Management explicitly identifies the gift as a scaling catalyst; later adoption and donors limit long-run attribution.
Gross modeled range: 1,000,000–3,000,000 incremental student-use episodes (central 2,000,000). After the separate impact-attribution range, Pine-attributable uplift is 100,000–900,000 (central 400,000). Calculation: Take 1m-3m of the post-grant adoption path as potentially enabled; assign 10%-30% counterfactual attribution to Pine. This yields 0.1m-0.9m uses, not learning gains.
Downstream and forecast update. A durable free writing platform and later AI product are plausible high-leverage effects. Upward on reach and product durability; still weak on causal learning magnitude.
Main caveat. User accounts and sentence counts are not unique learning outcomes.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. FY2018 spending was $7.067m; the system reported 176,201 people in its catchment and 160,521 patient visits. Pine equaled 14.1% of expense. Current work emphasizes evidence and government-owned models; no Pine-specific patient file exists. Large delivery share and organizational continuity; health outcomes are not grant-attributable.
Separate counterfactual estimate. I assign 40%–85% additional mission-resource share (central 65%), or 400, 000–850,000 of incremental mission resources (central $650,000). Material share of a resource-constrained system; bilateral/government funding and fixed facilities create displacement.
Gross modeled range: 40,000–70,000 patient-visit or person-year service equivalents (central 55,000). After the separate impact-attribution range, Pine-attributable uplift is 12,000–56,000 (central 30,250). Calculation: Contemporaneous cost benchmarks were about $14.76 per visit and $20.56 per capita; $1m gives 48.6k person-years or 67.8k visits. Use 40k-70k before attribution.
Downstream and forecast update. Possible helped transition models toward government and local partners; this may have durable system effects but is not quantified. Core care delivery occurred; later institutional transition makes naive lifetime patient attribution inappropriate.
Main caveat. Visits and catchment population are services, not DALYs; pre-grant mortality trends cannot be credited to Pine.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. The organization says the gift ‘breathed life’ into its work. Its 2018 Form 990 shows $75,665 of expense and $129,201 of revenue, so Pine exceeded one full year of spending. No family-level Pine output was located. Exceptionally material runway; service outcomes unresolved.
Separate counterfactual estimate. I assign 60%–95% additional mission-resource share (central 82%), or 60, 000–95,000 of incremental mission resources (central $82,000). Grant exceeded annual expense and recipient explicitly identifies an inflection, though some future donations may have been displaced.
Gross modeled range: 1–1.6 operating-year equivalents (central 1.3). After the separate impact-attribution range, Pine-attributable uplift is 0.6–1.5 (central 1.1). Calculation: $100k / $75,665 = 1.32 years at 2018 expense; apply counterfactual resource share.
Downstream and forecast update. Likely expanded resettlement volunteers, advocacy and household support, but no beneficiary denominator was published. Positive small-organization survival/capacity bet.
Main caveat. Runway is not the same as additional years of existence.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence D. The grant is publicly acknowledged as a $1m one-time gift. The Foundation continues clinical/research work and publications, but no spending schedule, patient count, or grant-specific research output was found. Receipt and organizational survival are verified; deployment remains unresolved.
Separate counterfactual estimate. I assign 35%–90% additional mission-resource share (central 65%), or 350, 000–900,000 of incremental mission resources (central $650,000). Large for a rare-disease nonprofit and likely highly useful; lack of finances and patient reporting widens downside.
No defensible outcome denominator; the estimate stops at additional mission resources. Calculation: Resource additionality only; no valid patient or DALY conversion.
Downstream and forecast update. Potential diagnostic and natural-history research value, unquantified. Unknown, with neither evidence of failure nor enough transparency to validate impact.
Main caveat. Current activity cannot be allocated to a single unrestricted gift.
Observed lookback — evidence B. Erowid reports $250k direct support and a $250k match that raised the required $250k in two months—about a normal year’s fundraising. It describes improved stability and project capacity; no user health outcome can be tied to the gift. Clear financial leverage and runway; harm-reduction outcome remains unmeasured.
Separate counterfactual estimate. I assign 55%–95% additional mission-resource share (central 80%), or 275, 000–475,000 of incremental mission resources (central $400,000). Gift and match were very large relative to normal fundraising; some matched donations may have shifted timing.
Gross modeled range: 0.8–1.8 operating-year equivalents (central 1.2). After the separate impact-attribution range, Pine-attributable uplift is 0.4–1.6 (central 0.8). Calculation: Recipient says the $250k match alone approximated ordinary annual fundraising; total Pine dollars likely represented roughly 1-2 years of prior operating scale.
Downstream and forecast update. Information access may avert risky drug use, but page views cannot be converted to health gains credibly. Financial/capacity success; endpoint impact still unobservable.
Main caveat. Matched dollars are not automatically new money; impact of information exposure is unmeasured.
Observed lookback — evidence A. The program was delayed about two years. A 2020 trial funded 12 projects with up to €50k; FY2021 accounts show £59,257 of grants against £186,790 of grant income, roughly one-third deployed by then. Restricted use was partly realized, but much more slowly and incompletely than the original gift implied.
Separate counterfactual estimate. I assign 35%–85% additional mission-resource share (central 60%), or 87, 500–212,500 of incremental mission resources (central $150,000). Restriction limits funging, but delay and partial deployment reduce realized incremental activity by the observed horizon.
Gross modeled range: 12–12 microgrants (central 12). After the separate impact-attribution range, Pine-attributable uplift is 9.6–12 (central 11.4). Calculation: Twelve grants are directly observed; the separate resource-share model discounts undeployed capital. No map-use benefit is inferred.
Downstream and forecast update. Community experimentation and local mapping capacity may persist; global OSM usage is not attributed. Downward on speed and absorption; positive on eventual pilot execution.
Main caveat. Currency conversion and later spending after 2021 are not fully reconciled.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. The gift was explicitly for long-term stability and supporting member projects, including crypto donation infrastructure. Against about $1.988m of annual expense it equaled 1.5 operating months. Financial capacity is clear; downstream software governance benefits are not allocable.
Separate counterfactual estimate. I assign 40%–85% additional mission-resource share (central 65%), or 100, 000–212,500 of incremental mission resources (central $162,500). Meaningful to a medium-sized organization with an explicit stability use; general support still funges.
Gross modeled range: 1.3–1.8 operating-month equivalents (central 1.5). After the separate impact-attribution range, Pine-attributable uplift is 0.5–1.5 (central 1). Calculation: $250k / $1.988m × 12 = 1.51 months; range covers expense-year variation.
Downstream and forecast update. Legal, fiscal-sponsorship, and compliance services protected multiple free-software projects, but user counts would be misleading. Likely modest positive resilience effect.
Main caveat. Runway equivalent does not imply counterfactual shutdown.
Observed lookback — evidence D. The organization said it had preserved at least ten brains and discussed storage costs around A$35k per case. No audited use-of-funds report, research output, or demonstrated benefit from the Pine gift was found. A small technical output is reported; grant deployment and scientific utility are unresolved.
Separate counterfactual estimate. I assign 15%–75% additional mission-resource share (central 40%), or 37, 500–187,500 of incremental mission resources (central $100,000). The grant was likely material, but transparency, throughput, and scientific validation are weak.
Gross modeled range: 4–10 brain-preservation equivalents (central 7). After the separate impact-attribution range, Pine-attributable uplift is 0.8–7.5 (central 3.2). Calculation: Use A$35k/case as an order-of-magnitude denominator and observed ten as a ceiling; apply low attribution.
Downstream and forecast update. Speculative archival option value only; no patient or consciousness outcome is claimed. Downward on transparency and demonstrated scientific value.
Main caveat. Preservation is not validated information recovery or future benefit.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. OMF says the gift transformed its research program: from about two projects per year to five in 2018 and around ten new projects per year in 2019-22, including a $1.8m commitment to a Harvard-affiliated center. No FDA-approved ME/CFS treatment has resulted by the cutoff. Large, documented capacity expansion; realized treatment DALYs remain zero/unknown.
Separate counterfactual estimate. I assign 55%–95% additional mission-resource share (central 80%), or 2, 750, 000–4,750,000 of incremental mission resources (central $4,000,000). Exceptionally large for OMF and explicitly described as transformational; later ME/CFS fundraising still contributes.
Gross modeled range: 8–25 research-project equivalents (central 15). After the separate impact-attribution range, Pine-attributable uplift is 2.8–21.3 (central 9). Calculation: Use the observed jump in portfolio size and OMF’s claim of >15 projects; estimate 8-25 project-equivalents before attribution.
Downstream and forecast update. Research networks and patient data may create option value. No DALYs are booked until a validated intervention changes care. Upward on scientific capacity, downward/deferral on cures and patient outcomes.
Main caveat. Projects differ radically in cost and value; counts are not scientific impact.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence D. The recipient says Pine created financial stability, kept the lights on, and allowed it to ramp up. No public counts of laboratories, reagent value, or shipments were found. A survival/capacity account exists; operational output is unresolved.
Separate counterfactual estimate. I assign 35%–90% additional mission-resource share (central 65%), or 35, 000–90,000 of incremental mission resources (central $65,000). Small, volunteer-heavy project says the grant was stabilizing; low public output and possible inactivity widen downside.
No defensible outcome denominator; the estimate stops at additional mission resources. Calculation: Resource additionality only; absence of shipment/value data blocks conversion.
Downstream and forecast update. Could have leveraged donated lab materials far above cash cost, but that leverage is not evidenced. Unresolved, with a meaningful chance of limited execution.
Main caveat. Original Pine ledger lists $100k despite a malformed transaction string; independent cash confirmation was not found.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. TalkLife grew from roughly 155k users in 2015 to claims of more than 5m people reached/member registrations. Later observational research found engagement associated with fewer self-injury thoughts but also risk from triggering content; an AI empathy experiment improved measured conversational empathy by 19.6%. A live, studied platform persists; the grant’s use and causal mental-health effect are unresolved.
Separate counterfactual estimate. I assign 20%–75% additional mission-resource share (central 45%), or 10, 000–37,500 of incremental mission resources (central $22,500). Small grant to a scalable product could matter, but commercial/nonprofit structure and subsequent capital are unclear.
Gross modeled range: 10,000–200,000 incremental user accounts (central 50,000). After the separate impact-attribution range, Pine-attributable uplift is 500–70,000 (central 7,500). Calculation: Assign a small slice of post-2018 platform growth to a $50k early grant; this is an explicitly speculative user-growth model, not health benefit.
Downstream and forecast update. Peer support and research infrastructure may have value; harms and substitution for professional care are not ruled out. Upward on platform survival/reach, ambiguous on net mental-health impact.
Main caveat. Registered/member counts, reach, and active users are different; no Pine use-of-funds report.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. Wings reports more than 100 missions, 1,800 flight hours, and 260k km since 2016. No Pine-funded aircraft, mission, or species outcome was identified. Meaningful activity followed the gift, but attribution is not documented.
Separate counterfactual estimate. I assign 30%–80% additional mission-resource share (central 55%), or 75, 000–200,000 of incremental mission resources (central $137,500). Early-stage organization and capital-intensive missions make the grant material; other conservation partners fund sorties.
Gross modeled range: 180–540 flight-hour equivalents (central 360). After the separate impact-attribution range, Pine-attributable uplift is 54–432 (central 198). Calculation: Take 10%-30% of 1,800 lifetime hours as the plausible Pine-enabled activity pool, then apply resource attribution.
Downstream and forecast update. Surveillance may protect habitat and wildlife, but no species or enforcement counterfactual is estimated. Plausibly positive implementation; ecological impact unresolved.
Main caveat. Flight hours are inputs, not hectares or animals protected.
Observed lookback — evidence B. The technology survives after merger. Current figures cite 1,800 researchers, 288 species, 300k animals reidentified, and 24 assessments. Microsoft AI for Earth support followed in 2018; no Pine-specific release or dataset was found. Durable technical platform and successor; grant contribution is not isolated.
Separate counterfactual estimate. I assign 35%–85% additional mission-resource share (central 60%), or 87, 500–212,500 of incremental mission resources (central $150,000). Early nonprofit funding likely bridged the team to later institutional support; Microsoft and others also contributed substantially.
Gross modeled range: 20,000–150,000 incremental animal reidentifications (central 60,000). After the separate impact-attribution range, Pine-attributable uplift is 3,000–90,000 (central 21,000). Calculation: Use 7%-50% of the 300k lifetime total as a Pine-enabled pool, then apply 15%-60% attribution. This measures database operations, not conservation outcomes.
Downstream and forecast update. Open-source identification tools plausibly improve research and population assessments; species protection effects remain unquantified. Upward on platform durability and adoption; unresolved on ecological outcomes.
Main caveat. An animal can be reidentified multiple times; this is not animals saved.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. Founded in 2018, Empower Work reports more than 432k workers ‘supported’ by 2024 across direct and content channels; among text-line users it reports 94% improved mental well-being and 77% economic improvement. Pine-specific contacts are not reported. Credible seed-stage survival and scale; reach definition and self-reported outcomes limit causal interpretation.
Separate counterfactual estimate. I assign 55%–95% additional mission-resource share (central 80%), or 55, 000–95,000 of incremental mission resources (central $80,000). Very early funding to a new service likely had high option and survival value; later donors drove scale.
Gross modeled range: 3,000–30,000 incremental direct support contacts (central 10,000). After the separate impact-attribution range, Pine-attributable uplift is 600–21,000 (central 4,500). Calculation: Exclude broad content reach; model 3k-30k direct-contact equivalents potentially enabled by seed capital, with 20%-70% attribution.
Downstream and forecast update. Volunteer training, employer partnerships, and career/income effects may persist; no causal wage or mental-health study was found. Upward on organizational scale; uncertain on measured welfare gains.
Main caveat. 432k includes broad reach and must not be treated as counseling episodes.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence A. ASF held the asset and recognized it in 2020. The realized gift equaled about 7.9 months of 2018 average spending or 6.4 months at 2020 spending. ASF already had 14.7 months of reserves in late 2017; it maintained hundreds of projects and near-continuous infrastructure. Clean balance-sheet resilience; no evidence of a counterfactual outage or project-level effect.
Separate counterfactual estimate. I assign 35%–80% additional mission-resource share (central 60%), or 350, 000–800,000 of incremental mission resources (central $600,000). Large unrestricted windfall with existing reserves and corporate sponsors; likely raised resilience more than current output.
Gross modeled range: 6.4–7.9 operating-month equivalents (central 7.1). After the separate impact-attribution range, Pine-attributable uplift is 2.2–6.3 (central 4.3). Calculation: 892, 882dividedbyreportedmonthlyspending(112.8k in 2018; $139.8k in 2020). Apply counterfactual resource share.
Downstream and forecast update. A defensible interpretation is 2-6 additional months of buffer/risk capacity, not a percentage of billions of software users. Smaller realized cash than headline and less dramatic operational need than rhetoric suggested; still useful resilience.
Main caveat. Reserve additions can affect risk-taking and future fundraising; months are not direct user benefit.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. The published plan proposed two FTEs and placing $750k in an investment portfolio. OpenMRS later reported 6,520 clinics/12.6m patients in 2021 and now claims more than 8,100 sites and 22m patients across 80+ countries, but those deployments are not Pine-attributable. Use-of-funds plan is specific; execution of the portfolio and hires is not fully reconciled publicly.
Separate counterfactual estimate. I assign 50%–90% additional mission-resource share (central 75%), or 500, 000–900,000 of incremental mission resources (central $750,000). Endowment structure creates durable incremental capacity and limits immediate funging; community development also had other funders.
Gross modeled range: 1–1 FTE-year plus endowment-capital package (central 1). After the separate impact-attribution range, Pine-attributable uplift is 0.5–0.9 (central 0.8). Calculation: Treat the published 250koperations/750k invested split as the unit; do not convert patient installations into health gains.
Downstream and forecast update. Durable open-source health-record infrastructure likely benefited many systems; incremental clinical effects remain unmeasured. Positive on platform reach and durability; still impossible to allocate patient outcomes.
Main caveat. A patient in an OpenMRS deployment is not necessarily a patient whose health improved because of the software.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. Its 2019-20 report records 111 people rescued, 573 supported, 117 arrests, 57 prosecutions and 17 convictions. The Victim Navigator program began in 2018 and later expanded from two to six police forces with markedly higher survivor engagement. No document assigns these outputs to Pine. Material post-grant programs and outputs; attribution remains unresolved.
Separate counterfactual estimate. I assign 35%–85% additional mission-resource share (central 60%), or 350, 000–850,000 of incremental mission resources (central $600,000). Large gift around launch of a scalable navigator model, but other grants, police partners and jurisdictions are essential.
Gross modeled range: 100–700 incremental survivor-support episodes (central 300). After the separate impact-attribution range, Pine-attributable uplift is 10–420 (central 90). Calculation: Use reported 2019-20 support and later navigator expansion to construct a 100-700 episode pool; apply low-to-moderate Pine attribution.
Downstream and forecast update. Training thousands of police/stakeholders may have institutional effects beyond cases; no welfare conversion is attempted. Plausibly positive implementation, with limited counterfactual evidence.
Main caveat. Rescues, support, arrests and convictions are distinct and should not be summed.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. Enthea published cost-effectiveness and policy analyses on psychedelics, addiction and drug liberalization. Its own model estimated a California ballot initiative at 26−92,950 per DALY, best guess $472, but no such campaign outcome is attributable to the grant. Intellectual outputs exist; no implemented policy or health outcome.
Separate counterfactual estimate. I assign 30%–85% additional mission-resource share (central 60%), or 15, 000–42,500 of incremental mission resources (central $30,000). Small research project and modest grant imply material support; unclear alternative funding and project completion.
Gross modeled range: 2–5 research-product equivalents (central 3). After the separate impact-attribution range, Pine-attributable uplift is 0.6–4.3 (central 1.8). Calculation: Count the public analyses rather than dividing $50k by the project’s own modeled $/DALY.
Downstream and forecast update. Potential influence on later advocacy is possible but not evidenced. Downward on implemented impact; reasonable on producing analysis.
Main caveat. Do not confuse this Enthea with the current psychedelic employee-benefit company; do not treat a modeled CEA as achieved DALYs.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. The grant supported planning for organ/biostasis research centers. By 2020 an NSF Engineering Research Center, ATP-Bio, received a roughly 26mfederalawardwithinabroader 80m partner commitment; in 2021 the Biostasis Research Institute launched with two centers. No transplant or patient outcome is Pine-attributable. Strong field-building sequence; very weak causal allocation to the large later consortium.
Separate counterfactual estimate. I assign 45%–90% additional mission-resource share (central 70%), or 900, 000–1,800,000 of incremental mission resources (central $1,400,000). Large early coordination grant plausibly helped a nascent field become fundable; federal and university partners dominate later scale.
Gross modeled range: 1–2 research-center equivalents (central 2). After the separate impact-attribution range, Pine-attributable uplift is 0.2–1.4 (central 0.9). Calculation: Attribute only part of the two-center successor infrastructure; separately report, but do not credit, the ~$80m ecosystem as leverage.
Downstream and forecast update. Potentially high option value for transplantation and preservation; zero realized patient DALYs are booked. Upward on institutional field-building, unresolved/downward on clinical translation.
Main caveat. The NSF award had many causal parents; assigning it to Pine would be severe overclaiming.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. FY2017-18 projected expense was $1.875m amid a revenue shortfall; Pine equaled 26.7% of spending. A 12-person staff supported more than 15k students and 718 instructors who added 11.5m words. The organization delayed vacancies and projects because revenue missed plan. Financially material at a constrained moment; output allocation is proportional rather than tagged.
Separate counterfactual estimate. I assign 55%–92% additional mission-resource share (central 78%), or 275, 000–460,000 of incremental mission resources (central $390,000). Unexpected gift offset a documented shortfall and delayed work, increasing additionality; later institutional grants still matter.
Gross modeled range: 2,500–6,000 student-year budget-equivalents (central 4,000). After the separate impact-attribution range, Pine-attributable uplift is 1,375–5,520 (central 3,120). Calculation: $500k / $1.875m × 15,000 = 4,000 students; similarly about 3.1m words. Range covers fixed costs and program mix.
Downstream and forecast update. The Dashboard and instructor network continued to grow; public knowledge readership is not attributed. More counterfactual than a generic big-organization grant because accounts show a binding shortfall.
Main caveat. Student editors, words, and article quality are different outcomes.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence U. The Australian practice remains active and says it has served thousands of clients over more than a decade. No Pine acknowledgment, grant use, subsidized-care count, or research output was found. Receipt is in Pine’s ledger; impact is opaque and charitable additionality is especially uncertain.
Separate counterfactual estimate. I assign 5%–60% additional mission-resource share (central 25%), or 2, 500–30,000 of incremental mission resources (central $12,500). For-profit/private-service structure raises displacement and private-benefit risk; a small early grant may still have financed access or product development.
Gross modeled range: 50–600 incremental client episodes (central 200). After the separate impact-attribution range, Pine-attributable uplift is 2.5–300 (central 40). Calculation: Speculative only: divide $50k by roughly 80−1,000 of marginal service/product cost, then heavily discount for commercial substitution.
Downstream and forecast update. Possible digital-therapy or practice growth, but no public causal chain. Low-transparency tail risk; one of the weaker lookbacks.
Main caveat. Do not equate current private clients with charitable beneficiaries.
Observed lookback — evidence B. Replate calls Pine one of its first major donors. It now reports 4.9m pounds recovered, 4.1m meals delivered, and 301 nonprofits served; its current donation benchmark is two rescued meals per dollar. No Pine-tagged meal count is published. Clear early capital and organizational survival; meal attribution is modeled, not observed.
Separate counterfactual estimate. I assign 50%–92% additional mission-resource share (central 75%), or 125, 000–230,000 of incremental mission resources (central $187,500). Early major donor in a logistics startup; later commercial partners and donations contribute to scale.
Gross modeled range: 300,000–700,000 meal equivalents (central 500,000). After the separate impact-attribution range, Pine-attributable uplift is 150,000–644,000 (central 375,000). Calculation: Current benchmark: $1 rescues two meals, so $250k = 500k; range allows historical costs and capacity spend.
Downstream and forecast update. Also avoided food waste and associated emissions/water use; these are not converted because attribution and lifecycle factors are uncertain. Strong positive on durable operations and visible output.
Main caveat. Meal-equivalent conventions do not guarantee meals consumed or nutritional quality.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence D. The Foundation continued research prizes, investments and field-building. No Pine-specific project or approved therapy was identified; many cited spinouts predate the grant or had other funders. Research option value may exist; grant-specific scientific output is unresolved.
Separate counterfactual estimate. I assign 30%–80% additional mission-resource share (central 55%), or 300, 000–800,000 of incremental mission resources (central $550,000). Material but fungible within a long-running foundation with multiple donors and investments.
No defensible outcome denominator; the estimate stops at additional mission resources. Calculation: Resource additionality only; no therapy probability model is defensible from public data.
Downstream and forecast update. Potential portfolio leverage is intentionally not summed because venture returns and scientific causality are opaque. Downward on realized health, unresolved on long-horizon option value.
Main caveat. Do not credit pre-2018 incubations or unrelated later mega-donations.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. The organization expanded to a fifth location and served 739 campers in 2019, 32% above the prior year (implying roughly 560 in 2018). No Pine-funded camper cohort or measured grief outcome was found. Expansion is visible and temporally close; grant attribution and clinical effect are unresolved.
Separate counterfactual estimate. I assign 45%–90% additional mission-resource share (central 70%), or 225, 000–450,000 of incremental mission resources (central $350,000). Large expansion gift for a growing nonprofit; other donors and volunteers share causality.
Gross modeled range: 250–500 camper-week equivalents (central 350). After the separate impact-attribution range, Pine-attributable uplift is 112.5–450 (central 245). Calculation: Use an order-of-magnitude 1k−2k fully loaded camp-week cost; $500k yields 250-500 before attribution.
Downstream and forecast update. New locations likely created recurring future capacity; no DALY or standardized grief reduction is claimed. Positive on reach and institutional durability.
Main caveat. Camp attendance is not a mental-health outcome; cost denominator is approximate.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. FSF publicly reported the conversion shortfall. Against contemporaneous annual expenses around 1.5m−1.8m, the realized gift equaled about 5.7-6.9 operating months. No campaign or software output was earmarked. Large runway with a documented volatility haircut; mission outcome not traceable.
Separate counterfactual estimate. I assign 35%–82% additional mission-resource share (central 60%), or 350, 000–820,000 of incremental mission resources (central $600,000). Very material relative to budget but unrestricted; membership revenue and later large gifts reduce additionality.
Gross modeled range: 5.7–6.9 operating-month equivalents (central 6.3). After the separate impact-attribution range, Pine-attributable uplift is 2–5.7 (central 3.8). Calculation: $860k / 1.5m−1.8m × 12.
Downstream and forecast update. Likely financed advocacy, licensing, GNU infrastructure and staff resilience; user freedom is not quantitatively allocated. Moderate positive runway, with a material BTC timing loss.
Main caveat. Do not confuse the Pine gift with a later separate $1m Handshake donation.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence D. The grant is confirmed in the Pine ledger, but no ACLU case, campaign, or allocation was found. Relative to the ACLU’s nine-figure post-2016 budget, it represented roughly days—not months—of national operations. Financial transfer is clear; impact is highly fungible and unresolved.
Separate counterfactual estimate. I assign 5%–50% additional mission-resource share (central 20%), or 100, 000–1,000,000 of incremental mission resources (central $400,000). Very large institution with abundant post-election fundraising implies high displacement; unrestricted cash may still fund marginal litigation or reserves.
No defensible outcome denominator; the estimate stops at additional mission resources. Calculation: Resource additionality only. Even an operating-day conversion would not identify cases or rights protected.
Downstream and forecast update. Possible litigation and policy value is real but cannot be separated from the portfolio. Likely lower marginality than grants to smaller organizations.
Main caveat. National and affiliate finances differ; no attempt is made to assign legal wins.
Observed lookback — evidence B. Women Who Tech identifies Pine among major supporters. Around the period, publicly named challenge prizes included about $35k plus a $25k Mozilla prize in 2018 and a $50k 2019 award. The broader network reports thousands of startups and substantial later fundraising, none fully Pine-attributable. Some direct prize outputs are visible; full $250k allocation and startup outcomes are incomplete.
Separate counterfactual estimate. I assign 40%–85% additional mission-resource share (central 65%), or 100, 000–212,500 of incremental mission resources (central $162,500). Program sponsorship likely expanded prize cohorts; corporate sponsors and judges also matter.
Gross modeled range: 4–10 startup-support packages (central 6). After the separate impact-attribution range, Pine-attributable uplift is 1.6–8.5 (central 3.9). Calculation: Use documented 25k−60k prize/support packages to estimate 4-10; do not credit subsequent venture capital raised.
Downstream and forecast update. Founder survival and financing may have large upside, but return data are selected and not counterfactual. Moderately positive direct support; long-run entrepreneurial impact uncertain.
Main caveat. Cohort fundraising rates have survivorship and selection bias.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence D. Blueprint teams of about five students build year-long pro bono software projects for nonprofits. The network now has multiple chapters and project archives, but no Pine-funded team, partner, or codebase was identified. Organization persists; grant deployment is unresolved.
Separate counterfactual estimate. I assign 35%–88% additional mission-resource share (central 65%), or 17, 500–44,000 of incremental mission resources (central $32,500). Very small volunteer organization likely found $50k material; university ecosystem and volunteer labor are alternative inputs.
Gross modeled range: 2–8 year-long software-project equivalents (central 4). After the separate impact-attribution range, Pine-attributable uplift is 0.7–7 (central 2.6). Calculation: Order-of-magnitude 6k−25k cash support per volunteer software project, excluding donated labor.
Downstream and forecast update. Beneficiary nonprofits may retain software and student alumni may carry skills forward; neither is quantified. Plausible capacity success with weak reporting.
Main caveat. Volunteer labor value can dwarf cash and should not be attributed to Pine.
Observed lookback — evidence A. The largest gift in the organization’s history financed a building/down payment and scaling. 2018 expense was $331k, so Pine equaled 75% of annual spending. The program now reports about 90 therapy teams, 1,400 hours and more than 5,000 people served annually. Concrete durable asset and scale are well supported; patient effects are not separately measured.
Separate counterfactual estimate. I assign 65%–97% additional mission-resource share (central 85%), or 162, 500–242,500 of incremental mission resources (central $212,500). Largest-ever gift and designated capital use make it highly additional; another donor or lease was still possible.
Gross modeled range: 250–600 annual therapy-hour capacity added (central 400). After the separate impact-attribution range, Pine-attributable uplift is 112.5–540 (central 280). Calculation: Compare current ~90 teams/~1,400 hours with roughly 65 teams around the grant; infer 250-600 annual hours of added capacity, then discount attribution.
Downstream and forecast update. A building can generate many years of service and fundraising stability; no QALY/DALY conversion is warranted. One of the clearest durable-capital successes.
Main caveat. Team and hour comparisons are not a controlled pre/post series.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence D. IJ said the gift would help defend rights of hundreds of thousands but did not earmark it to cases. Litigation, legislative and communications activity continued; no defensible Pine case count was found. Mission use is plausible; outcome attribution is unresolved.
Separate counterfactual estimate. I assign 15%–65% additional mission-resource share (central 35%), or 300, 000–1,300,000 of incremental mission resources (central $700,000). Established national organization and fungible legal portfolio create displacement; $2m is still material for riskier cases.
No defensible outcome denominator; the estimate stops at additional mission resources. Calculation: Resource additionality only; cases differ too much in cost and population effect for a count.
Downstream and forecast update. Precedent and policy effects may be large and asymmetric, but a probability-weighted legal model would be mostly invented. Unresolved; likely less marginal than small-grantee grants.
Main caveat. IJ’s ‘hundreds of thousands’ statement is prospective rhetoric, not a retrospective count.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. TreeSisters reports 28.68m trees funded to date and a 50% female workforce across projects. No Pine-tagged cohort or survival monitoring was located. Large planting program continued; grant-specific and survival outcomes are unresolved.
Separate counterfactual estimate. I assign 50%–92% additional mission-resource share (central 75%), or 250, 000–460,000 of incremental mission resources (central $375,000). Material growth capital in a then-young network; fundraising substitution and tree mortality reduce attribution.
Gross modeled range: 500,000–1,500,000 trees funded (central 1,000,000). After the separate impact-attribution range, Pine-attributable uplift is 250,000–1,380,000 (central 750,000). Calculation: Historic public messaging put planting around $0.50/tree; range 0.33−1.00 per tree. This counts funded seedlings, not mature surviving trees or carbon.
Downstream and forecast update. Ecological and livelihood effects may last decades; no tCO2e is assigned without species, location, survival and leakage data. Positive on scale, still uncertain on durable ecological outcome.
Main caveat. Trees funded ≠ trees planted ≠ trees surviving ≠ additional carbon removal.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. The recipient announced expansion of device, connectivity and training work. It now reports more than 24k free devices annually; no Pine-tagged recipient list was found. Use category and later scale are clear; exact output attribution is not.
Separate counterfactual estimate. I assign 45%–90% additional mission-resource share (central 70%), or 112, 500–225,000 of incremental mission resources (central $175,000). Material grant to a scaling refurbisher; donated hardware, contracts and later funders are major inputs.
Gross modeled range: 1,000–3,000 device/household equivalents (central 1,900). After the separate impact-attribution range, Pine-attributable uplift is 450–2,700 (central 1,330). Calculation: Using roughly 80−250 marginal cash cost per refurbished device/connected household; a public low-cost laptop benchmark near $130 gives ~1,923.
Downstream and forecast update. Employment, education and connectivity gains likely extend beyond devices; no income or learning conversion is made. Positive on continuing scale and clear service type.
Main caveat. Devices may be free, subsidized or program-funded; recipient overlap and durability are unknown.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. The organization says Pine enabled staff growth from two to four, stronger systems, and a policy to prioritize at least 50% women/local founders. It now reports 2,747 jobs and 2.139m people reached; a 2020 snapshot showed 315 jobs with $340,473 in grants. Clear organizational uplift and policy change; job and beneficiary counts are not grant-tagged.
Separate counterfactual estimate. I assign 55%–95% additional mission-resource share (central 80%), or 550, 000–950,000 of incremental mission resources (central $800,000). Very large relative to a two-person staff and recipient explicitly credits expansion; partner financing and later donors still drive jobs.
Gross modeled range: 500–1,500 livelihood-job equivalents (central 900). After the separate impact-attribution range, Pine-attributable uplift is 150–1,200 (central 495). Calculation: 2020 grants imply ~$1,081 per job; $1m ≈925. Range 500-1,500 for year/program variation, then discount for capacity use and displacement.
Downstream and forecast update. Each trained worker is claimed to serve hundreds of neighbors, but those indirect reach multipliers are not converted to welfare. Strong positive organizational lookback; beneficiary effects need better causal evidence.
Main caveat. Jobs created and people served are partner-reported; job duration and income gains are not standardized.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence A. Contemporaneous reports state Pine launched the accelerator. SOA later reported $4.1m invested in startups across 22 countries, $1.6m in microgrants to 300 grantees in 84 countries, and portfolio companies raising hundreds of millions; no environmental outcome ledger exists. Foundational program launch is clear; startup success and ocean outcomes have many causes.
Separate counterfactual estimate. I assign 65%–97% additional mission-resource share (central 85%), or 650, 000–970,000 of incremental mission resources (central $850,000). The $1m was explicitly launch capital, making it highly counterfactual; later partners, founders and investors dominate downstream scale.
Gross modeled range: 20–120 startups/grantees materially enabled (central 50). After the separate impact-attribution range, Pine-attributable uplift is 3–72 (central 17.5). Calculation: Use 300 grantees plus accelerator cohorts as the upper pool and assign a small early-funder share; do not count follow-on capital as impact.
Downstream and forecast update. Network firms reportedly raised 225m−550m and created 950+ jobs; Pine’s causal share might be 5%-20%, but this is leverage, not ocean benefit. Upward on ecosystem formation and capital mobilization; unknown on net environmental outcomes.
Main caveat. Capital raised is not social value and may include commercial returns.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. The 2018 report shows $1.65m raised, $257k caring for 60 girls, $600k completing a 900-place school, $314,781 operating school services for 90, and $45,882 for advocacy. Pine was about 15% of revenue; no line-item assignment is made. Contemporaneous program accounts permit a useful denominator; allocation remains proportional.
Separate counterfactual estimate. I assign 45%–90% additional mission-resource share (central 70%), or 112, 500–225,000 of incremental mission resources (central $175,000). Material to a small program with concrete school/care expenses; construction and fundraising substitution remain.
Gross modeled range: 40–90 child-year care/education equivalents (central 65). After the separate impact-attribution range, Pine-attributable uplift is 18–81 (central 45.5). Calculation: 250k/(257k/60) = 58 care-years; or /($314,781/90)=71 school-years. Range adds program mix.
Downstream and forecast update. A 900-place school creates capacity beyond one year, but occupancy and learning are not attributed. Likely delivered substantial care/education capacity as expected.
Main caveat. Care-years and school-years are different services and appear only as a bracket.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. Against a roughly $3m 2018 budget, Pine represented about one month of operations. Let’s Encrypt scaled from about 1m certificates/day in 2018 to around 10m/day and near one billion active certificates by 2025; no part of that reach can be isolated to Pine. A useful but small slice of a rapidly scaling infrastructure budget.
Separate counterfactual estimate. I assign 25%–75% additional mission-resource share (central 50%), or 62, 500–187,500 of incremental mission resources (central $125,000). Material at the time but supported by major corporate sponsors; donation likely improved buffer more than determined scale.
Gross modeled range: 0.8–1.2 operating-month equivalents (central 1). After the separate impact-attribution range, Pine-attributable uplift is 0.2–0.9 (central 0.5). Calculation: $250k / $3m × 12 = 1 month; apply counterfactual resource share.
Downstream and forecast update. TLS adoption created immense security/privacy value, but multiplying certificates by Pine’s budget share would be spurious. Upward on platform scale, modest on Pine’s marginal share.
Main caveat. Certificates, domains and people are not interchangeable.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence D. Snow Lovers says it joined/started in 2018 because of the Pine funding and presents climate and diversity education. No participant, curriculum reach, financial or outcome report was found. Organizational creation may be attributable; execution beyond a website is opaque.
Separate counterfactual estimate. I assign 20%–85% additional mission-resource share (central 50%), or 10, 000–42,500 of incremental mission resources (central $25,000). Startup capital can be decisive, but public evidence of activity is unusually thin.
No defensible outcome denominator; the estimate stops at additional mission resources. Calculation: Resource additionality only; no credible participant denominator.
Downstream and forecast update. Possible community and awareness value, unquantified. High risk of limited realized impact; one of the least observable grants.
Main caveat. Website continuity is weak evidence of program delivery.
Observed lookback — evidence B. The recipient said Pine enabled rapid expansion to about 40 chapters on five continents; it later described 50 chapters on six continents and launched 5k−10k Black Founder Startup Grants. No reconciliation of the $1m or complete grantee list was found. Clear network-scale inflection; founder economic outcomes remain unmeasured.
Separate counterfactual estimate. I assign 55%–95% additional mission-resource share (central 80%), or 550, 000–950,000 of incremental mission resources (central $800,000). Recipient explicitly attributes rapid chapter expansion to a very large gift; chapters and later sponsors supply labor/capital.
Gross modeled range: 30–100 chapter or founder-grant equivalents (central 50). After the separate impact-attribution range, Pine-attributable uplift is 10.5–85 (central 30). Calculation: Treat 30-50 chapters plus a possible pool of 5k−10k grants as comparable support packages only for order of magnitude; do not add both.
Downstream and forecast update. Founder networks may improve capital access; later venture success is highly selected and not attributed. Upward on global network creation; unclear on durable founder income/firm survival.
Main caveat. Chapters vary in activity; the current grant program has other sponsors.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence A. TRGT reports that six parcels totaling 409 newly protected acres connected 3,157 acres across more than seven miles; Pine was among the funders. The purchase stack and Pine’s share are not published. A concrete acquisition is tied to the funding cohort, but not apportioned among donors.
Separate counterfactual estimate. I assign 50%–92% additional mission-resource share (central 75%), or 125, 000–230,000 of incremental mission resources (central $187,500). Restricted acquisition capital is additional, though other donors might have closed the gap or delayed purchase.
Gross modeled range: 50–300 acres protected (central 150). After the separate impact-attribution range, Pine-attributable uplift is 20–270 (central 97.5). Calculation: Use the 409-acre joint acquisition as the ceiling and assign 12%-73% based on plausible financing share; then apply counterfactual closure probability.
Downstream and forecast update. Connectivity may increase ecological value beyond acres; no biodiversity or carbon conversion is attempted. Strong concrete conservation output with allocation uncertainty.
Main caveat. The 3,157-acre connected landscape is not 3,157 newly acquired Pine acres.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. KDE acknowledged €200k-equivalent support and continued funding developer sprints, conferences and infrastructure. Relative to historical annual income around $126k in 2015, the gift exceeded one prior-year revenue; no expense ledger ties outputs to Pine. Exceptionally material to the nonprofit shell; software/user outcomes remain diffuse.
Separate counterfactual estimate. I assign 50%–92% additional mission-resource share (central 75%), or 100, 000–184,000 of incremental mission resources (central $150,000). Large relative to prior income and volunteer ecosystem; corporate/community donations create some substitution.
Gross modeled range: 1.3–2 historical annual-income equivalents (central 1.6). After the separate impact-attribution range, Pine-attributable uplift is 0.7–1.8 (central 1.2). Calculation: $200k divided by roughly 100k−150k historical annual income.
Downstream and forecast update. Likely financed contributor coordination and infrastructure. Desktop users are not used as an impact multiplier. Strong resource uplift; end-user benefit unresolved.
Main caveat. Revenue-equivalent is not an operating-year estimate and uses an older benchmark.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence A. The grant paid off the ranch mortgage, a concrete durable asset that freed later donations for care. Nuzzles now reports more than 27,000 animals saved over 35 years; Pine-specific additional adoptions are not published. One of the clearest capital lookbacks; animal outcomes require a facility-capacity model.
Separate counterfactual estimate. I assign 75%–100% additional mission-resource share (central 90%), or 750, 000–1,000,000 of incremental mission resources (central $900,000). Exact asset use and debt retirement make displacement low, though another capital campaign was possible.
Gross modeled range: 300–1,000 annual animal-care capacity equivalents (central 600). After the separate impact-attribution range, Pine-attributable uplift is 90–800 (central 330). Calculation: Assume the debt-free ranch enabled 300-1,000 animal outcomes annually for several years; apply a modest Pine share because staff, adopters and operating donors are necessary.
Downstream and forecast update. The durable balance-sheet benefit could persist for decades; counting all later rescues would overstate causality. Strong positive and unusually concrete.
Main caveat. Animals ‘saved’ is organization-defined; the annual capacity estimate is speculative.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence B. The Pine gift arrived after GiveWell had granted $5.944m intended to fund the vaccination program/RCT through May 2020. The RCT later found a 22 percentage-point increase in full immunization. At mature scale GiveWell’s 2025 lookback found about $18 per enrolled child and $2,117 per death averted, roughly twice its initial estimate. The intervention became exceptionally effective, but Pine’s marginal 2018 grant may have funded runway/RCT support rather than additional vaccinations.
Separate counterfactual estimate. I assign 20%–80% additional mission-resource share (central 50%), or 50, 000–200,000 of incremental mission resources (central $125,000). Program quality is high, but a prior GiveWell grant purported to fully fund the RCT period, creating unusually important funging/timing uncertainty.
Gross modeled range: 42–167 mature-program deaths-averted benchmark (central 118). After the separate impact-attribution range, Pine-attributable uplift is 8.4–133.6 (central 59). Calculation: $250k divided by 1, 500−6,000 per death (central $2,117). Apply a separate 20%-80% grant additionality range because 2018 deployment is not known. At 50-70 DALYs/death, the gross benchmark is roughly 2.1k-11.7k DALYs, not a lookback finding.
Downstream and forecast update. Early unrestricted runway may have helped the organization reach the 2020 top-charity/scale inflection, potentially making true downstream impact much larger than direct deployment—but attribution is highly uncertain. Strong upward intervention update; grant-specific counterfactuality is lower than a naive mature-cost calculation suggests.
Main caveat. The 2021-23 mature program’s costs and mortality model are not 2018 realized outcomes.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. 2018 revenue was $697,771 and expense $423,642; Pine was 7.2% of revenue or 1.4 months of spending. MEAction continued ME/CFS advocacy and later Long COVID work; no campaign or policy win was tied to the gift. Moderate runway; advocacy impact unresolved.
Separate counterfactual estimate. I assign 35%–82% additional mission-resource share (central 60%), or 17, 500–41,000 of incremental mission resources (central $30,000). Material to a young movement organization but fungible; later disease attention and donors are major drivers.
Gross modeled range: 1.1–1.8 operating-month equivalents (central 1.4). After the separate impact-attribution range, Pine-attributable uplift is 0.4–1.5 (central 0.8). Calculation: $50k / $423,642 × 12 = 1.42 months; range allows expense timing.
Downstream and forecast update. Potential research funding, recognition, and patient-support effects are not converted to DALYs. Positive capacity, unquantified advocacy effect.
Main caveat. Long COVID work largely postdates the gift and must not be automatically credited.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence A. The Religious Leaders Study’s final paper appeared in 2025. Thirty-three clergy were enrolled; most reported enduring spiritual significance, and no serious adverse events were reported. The sample was small and homogeneous; Johns Hopkins found serious IRB noncompliance including incomplete funding disclosures. A completed, published study is a clear output; welfare impact and research quality are much more ambiguous.
Separate counterfactual estimate. I assign 45%–90% additional mission-resource share (central 70%), or 112, 500–225,000 of incremental mission resources (central $175,000). Grant size plausibly covered a large study share; university resources and other funders also mattered.
Gross modeled range: 24–33 trial-participant experiences (central 33). After the separate impact-attribution range, Pine-attributable uplift is 9.6–31.3 (central 23.1). Calculation: Use enrolled/analyzed participant counts and apply study-financing attribution. No DALYs or spiritual-value unit is imposed.
Downstream and forecast update. May affect scholarship and religious discourse; policy/clinical effects are not demonstrated. Research output realized after a long delay, with governance concerns lowering confidence.
Main caveat. Self-reported spiritual significance is not clinical benefit; compliance concerns matter.
Sources: source 1 · source 2 · source 3.
Observed lookback — evidence C. The nonprofit had 2017 revenue of $201,920, expense $302,799 and negative net assets. In 2018 revenue jumped to $4.15m, so Pine was only about 24% of that year’s inflow; by 2024 filings showed about $808k net assets after investment sales. No beneficiary or portfolio-impact report isolates Pine. The grant materially recapitalized a thin nonprofit, but other 2018 funding was larger and social returns are opaque.
Separate counterfactual estimate. I assign 20%–75% additional mission-resource share (central 45%), or 200, 000–750,000 of incremental mission resources (central $450,000). Severe prior capital constraint raises marginality, but $3m+ other 2018 revenue and investment asset sales reduce Pine’s unique role.
No defensible outcome denominator; the estimate stops at additional mission resources. Calculation: Resource additionality only; a venture/job/borrower count would require portfolio-level data that are not public.
Downstream and forecast update. Investment stakes in firms such as Pathao and CreditEnable may have created services and returns, but neither commercial value nor social impact is attributed. More ambiguous than the 2018 ‘high-impact entrepreneurship’ story; transparency is weak.
Main caveat. Do not assume the current VC firm’s activities or returns belong to the original nonprofit grant.
Sources: source 1 · source 2 · source 3.
The accompanying workbook contains the 60-row ledger, evidence grades, a separate modeled-attribution sheet, formula-backed unit calculations, assumptions, and one-URL-per-cell source index. The most important portfolio sources are the original Pine ledger, original announcement, dated cross-check ledger, and recipient follow-up interviews. GiveDirectly and New Incentives calculations use GiveWell’s current technical pages and corrected lookbacks, linked in their grant entries.
This report intentionally leaves unlike outcomes unlike. A later decision model could add explicit moral weights, discounting, and risk preferences, but doing so here would conceal rather than resolve the normative choices.
Broad bucket | Nominal amount | Share |
|---|---|---|
Medical & scientific research | $16,650,000 | 30% |
Poverty, livelihoods & inclusion | $9,700,000 | 17% |
Rights, justice & equity | $7,250,000 | 13% |
Digital & public infrastructure | $7,050,000 | 13% |
Education | $4,000,000 | 7% |
Global health delivery | $3,250,000 | 6% |
WASH | $3,000,000 | 5% |
Environment & conservation | $2,400,000 | 4% |
Animal welfare | $1,250,000 | 2% |
Health support & advocacy | $1,200,000 | 2% |
Some notes by me based on Sol’s writeup:
The ~$86M nominal value at the outset ended up being worth $55.75M due to BTC volatility and something like $39M (range $26-50M) after counterfactual adjustments (gift restriction, grantee size, evidence of a binding funding constraint, absorption/delay, other funders)
Health-wise: ~120 lives saved via New Incentives, ~45k annual income doubling equivalents via GiveDirectly, ~3,200 patients treated via Watsi, ~40,000 people get clean drinking water, ~70,000 sick people got hospital visits via Possible Health
I’m kind of bummed the MAPS donation hasn’t yet materialised endline benefits
Maybe 1-7 brains counterfactually preserved by the Neural Archives Foundation
A bunch others, tremendous variety, check out the writeups above if you’re interested
ACX Grants 1-3 year updates for contrast ($3M)
The first cohort of ACX Grants was announced in late 2021, the second in early 2024. In 2022, I posted one-year updates for the first cohort. Now, as I start thinking about a third round, I’ve collected one-year updates on the second and three-year updates on the first. …
The total cost of ACX Grants, both rounds, was about $3 million. Do these outcomes represent a successful use of that amount of money? …
It’s harder to produce Inside View estimates, because so many of the projects either produce vague deliverables (eg a white paper that might guide future action) or intermediate results only (eg getting a government to pass AI safety regulations is good, but can’t be considered an end result unless those regulations prevent the AI apocalypse). Because we tend towards incubating charities and funding research (rather than last-mile causes like buying bednets), achieved measurable deliverables are thin on the ground. But here are things that ACX grantees have already accomplished:
Improved the living/slaughter conditions of 30 million fish.
Helped create Manifold Markets, a prediction market site with thousands of satisfied users, whose various spinoffs play a central role in the rationalist/EA community.
Helped create thousands of jobs in Rwanda and other developing countries
Passed an instant runoff vote proposition in Seattle.
Saved between a few dozen and a few hundred lives in Nigeria through better obstetric care.
And here are some intermediate deliverables from grantees:
Made Australian government take AI x-risk more seriously (estimated from 50th percentile to 60th percentile outcome)
Gotten the End Kidney Deaths Act (could save >1000 lives and billions of dollars per year) in front of Congress, with decent odds of passing by 2026.
Plausibly saved 2 billion chickens from painful death over next decade2.
Antiparasitic medication oxfendazole continues to advance through the clinical trial process.
And here are some things that have not been delivered yet but that I remain especially optimistic about:
Creation of anti-mosquito drones that provide a second level of defense along with bednets.
Revolutionize diagnosis of traumatic brain injury
Improve dietary guidelines in developing countries
Continue to support research and adoption of far UV light for pandemic prevention
Reduce lead poisoning in Nigeria
I think these underestimate success since many projects have yet to pay off (or to convince me to be especially optimistic), and others have paid off in vague hard-to-measure ways.
I remain impressed from a purely grantmaking logistics perspective that “just Pine and a friend” were able to go through 10,000+ applications and donate $50M+ to 60 orgs in just 4 months, and that the grantee selection seemed not half bad at all, all while remaining anonymous. There’s something aesthetically pleasing about this.
Out of curiosity I asked Sol (max) what Screwtape’s Mythos-powered Dr Doom might be able to do:
Imagine the date is August 15th, 2016. You, and only you, have access to Mythos. Everyone else in the world has only the machine learning tools of 2016. That’s a crazy advantage, right? That is the kind of supertech that comic books are made of. You’d look like Dr Doom.
Sol’s overall take:
The most impressive capability-only achievement would be building, several years early, a tiny institute that continuously turns public code and public data into validated security fixes, decision products and occasional research discoveries. Its apparent superpower would come from concentrating vast machine search behind unusually hard feedback loops. Outside those loops, Dr Doom would still spend a depressing amount of time waiting for customers, maintainers, lawyers and laboratory results. …
Sol’s full answer
The counterfactual contains a loophole large enough to dominate the engineering question. A 2026 model transported intact to August 2016 is a lossy compressed record of the following ten years. It remembers asset prices, company outcomes, elections, vulnerabilities, scientific discoveries, successful products, COVID-19, and the approaches which eventually worked on thousands of then-open problems.
That produces two very different answers:
In the literal counterfactual, you possess an unreliable time oracle. The highest-impact use is to calibrate that oracle, quietly accumulate capital, and use the capital plus foreknowledge to build institutions or mitigate harms.
In the more interesting capabilities-only counterfactual, with post-August-2016 facts quarantined, you can probably create a tiny, self-financing software, data, security, and research institute. It would be extraordinarily productive wherever reality supplies cheap, hard feedback. It would remain surprisingly ordinary at sales, taste, wet-lab work, institutional trust, and deciding what actually matters.
The second case is harder and more revealing, so most of my answer concerns it.
“Give me Mythos in 2016” sounds like transporting a capability. In fact, you transport capabilities, factual memory, later algorithms, later software conventions, and ten years of examples of what succeeded.
There is direct evidence that historical-agent benchmarks already encounter this. In METR’s NanoGPT experiments, models sometimes proposed improvements resembling later leaderboard contributions; one Opus run explicitly identified future contributions. METR had to distinguish genuine search from knowledge of the benchmark’s subsequent history in its expenditure-horizon analysis. Historical Kaggle contests, old repositories, famous conjectures, and retrospective vulnerability searches are all vulnerable to the same contamination.
I would therefore state three versions of the thought experiment:
Version | What crosses into 2016 | Consequence |
|---|---|---|
Literal | 2026 weights, memories, tools and inference | A noisy ten-year oracle dominates |
Capability-only | 2026 reasoning and agency, with post-2016 factual answers forbidden | Interesting, although the firewall leaks through algorithms and priors |
Truly trained only on pre-2016 material | Model architecture and training occur using only 2016 data | Much less capable; many impressive present-day results disappear |
The literal version should begin with a prediction ledger. Ask several model families separately for hundreds of dated, objectively resolvable claims over the next three, six, and twelve months. Record probabilities, hash or notarize the predictions, and refrain from intervening in the relevant events. Consensus between models is only weak evidence because their training corpora overlap. After enough predictions resolve, estimate calibration by topic and model.
Once that works, diversified lawful investment is the obvious funding mechanism. It is less operationally demanding than building twenty businesses and initially less likely to alter the events being predicted. As interventions accumulate, the remembered future becomes progressively less reliable because the timeline diverges.
The most impressive ethical project in this version is probably neither an engineering company nor a bug-bounty career. It is:
accumulate a great deal of capital without advertising the source of the advantage;
build broad pandemic preparedness, genomic surveillance, vaccine-platform capacity, PPE manufacturing, and public-health logistics before 2020;
privately alert vendors to well-verified future vulnerabilities through responsible disclosure;
fund scientific projects whose future importance the models remember;
preserve a dated forecast archive so that future knowledge can be distinguished from later model improvisation.
I would avoid publishing pathogen-specific information or attempting to reproduce particular future biological events. The safe edge lies in broad preparedness and platform capacity. Much of the model’s detailed future memory would become stale after significant interventions anyway.
There is another hidden assumption. If only open weights cross the time boundary, 2016 hardware cannot economically run the largest 2026 models. The premise’s $5,000 API budget implicitly gives you a magical cross-temporal inference service. I would also assume that the service can use a contemporary sandbox but has no live 2026 internet connection. If it has 2026 web access, the oracle interpretation becomes even more overwhelming.
The most misleading way to reason here is to treat “Mythos,” “Astra,” or “Sol Ultra” as scalar quantities of intelligence. The impressive results are model-plus-campaign systems: persistent state, many rollouts, tailored tools, executable feedback, human problem selection, and some means of rejecting plausible nonsense.
A few current results show the shape of that dependence unusually clearly:
Result | What happened | Hidden structure and 2016 lesson |
|---|---|---|
Anthropic’s C compiler | Sixteen agents produced roughly 100,000 lines of Rust over two weeks for about $20,000. It compiled Linux for several architectures and built software including PostgreSQL, SQLite, Redis, QEMU and FFmpeg. | GCC served as a differential oracle. The result still relied on GCC for part of 16-bit x86, lacked a stable independent assembler/linker, emitted worse code than GCC |
OpenAI’s internal agent-built product | Three engineers oversaw roughly one million lines of code and 1,500 pull requests over five months, with no human-written application code. | Humans made the repository, specifications, logs, metrics and user interface legible to the agents. Human QA remained a bottleneck. OpenAI |
Anthropic recursive-improvement experiments | About 800 agent-hours and $18,000 recovered 97% of a benchmark gap; another campaign made hundreds of fixes and greatly reduced API errors. | Humans chose the objective and scoring. Some improvements failed to transfer to production. Anthropic’s staff reported large gains, but the company explicitly warns that lines of code and self-estimated speedups overstate value. Anthropic |
Remote Labor Index | The latest reported leader fully automated 15.8% of projects. Earlier top systems were in single digits. Projects can exceed 100 hours and $10,000. | Current systems still fail professional acceptance on a large majority of end-to-end commissioned work. The gap between a runnable artifact and accepted professional work is substantial. RLI, CAIS update |
GDPval | GPT-5.2 beat or tied experts on 70.9% of well-specified knowledge-work tasks, at over eleven times the speed and below one percent of expert cost. | This sits beside RLI rather than contradicting it. Specification and acceptance conditions determine much of the observed capability. OpenAI |
Experienced-developer RCT | Sixteen developers completed 246 tasks on familiar open-source repositories and were 19% slower using early-2025 AI, despite predicting they would be faster and afterward believing they had been faster. | Familiar-codebase context, review, and correction costs can cancel generation speed. Later METR work found that task avoidance and parallel agents make the newer effect difficult to estimate cleanly. Initial study, 2026 update |
NVIDIA’s Kaggle system | Three agents conducted roughly 850 experiments, produced a 600,000-line repository and assembled a 150-model stack that won a competition. | A leaderboard supplied a dense objective and made enormous search useful. NVIDIA |
The shared pattern is more useful than any one number. Models are strongest when the world can be turned into an executable specification: tests, compilers, simulators, formal kernels, leaderboards, differential oracles, or rapid customer behavior. They are much weaker when progress is judged through tacit professional standards, delayed physical evidence, confused customers, politics, or aesthetics.
Even on NanoGPT, where the objective is unusually clean, METR found that raw gains were inflated by noisy measurements, brittle optimizations and reward hacking. After revalidation, spending above $10,000 bought roughly another 1 to 1.5 percent in the strongest campaigns. Many model-generated ideas were reasonable enough to merge, while a smaller subset generated actual speedup. That is a good picture of your 2016 advantage: a vast supply of plausible work, with value concentrated in the branches that reality can cheaply reject.
Your compilation of high-token-use cases is directionally persuasive, but “tokens per month” obscures much of the economics.
SemiAnalysis reports about five billion tokens per employee per month, with some users exceeding 100 billion. Its characteristic workflow has a roughly 300:1 input-to-output ratio, more than 90% cached input, and an estimated blended cost near $0.99 per million tokens. It also says model expenditure reaches about 30% of compensation in some workflows. Those are highly optimized context-reuse economics, rather than five billion freshly reasoned output tokens. See the underlying AI value-capture analysis and its interviews about token budgeting.
At that blend, $5,000 corresponds to about five billion token-events. At Mythos’s current listed rates of $10 per million input and $50 per million output, a rollout-heavy campaign can consume the same budget vastly faster. A 100-billion-token user is usually exploiting caching, routing, reused repositories and cheap scout models. Comparing that user to an Astra research campaign by raw tokens is nearly meaningless.
Some concrete cost anchors are more informative:
Campaign | Reported or inferable model cost | Initial-budget equivalent |
|---|---|---|
Astra’s ten mathematical advances | About $2,000 in discovery tokens at Sol API rates, excluding much human manuscript work | 0.4 month |
Anthropic Riemann campaign | 31 million output tokens; at $50 per million, output alone would be $1,550, with input, tools and internal-model pricing additional | At least 0.3 month |
Anthropic C compiler | About $20,000 | 4 months |
Mythos OpenBSD campaign | Fewer than $20,000 for about 1,000 runs | Under 4 months |
Anthropic recursive benchmark campaign | $18,000 | 3.6 months |
Smart-contract scan | $3,476 | 0.7 month |
These campaigns differ so much that the table should not be used to fit a scaling law. It establishes a narrower point: $5,000 per month can fund serious searches, although one large campaign can absorb several months. The initial user cannot continuously run C-compiler-scale, Mythos-scale, Riemann-scale and product-development-scale efforts in parallel.
The first compounding advantage is therefore campaign engineering. Cheap models scout, extract and propose. Strong coding models implement. Mythos is reserved for explicitly authorized defensive-security targets. Astra and the internal mathematics models receive problems with unusually good verification surfaces. The strongest general model writes specifications, audits plans and adjudicates disagreements. Expensive multi-agent swarms are deployed only when branches are genuinely separable or a hard external verifier exists.
The prolific power users in your collection also tend to be experienced operators. The Liu Xiaopai and Jason Hoffman stories are evidence about what a strong model does when attached to decades of accumulated product judgment, system architecture and distribution knowledge. They provide weaker evidence about what a random clever person could obtain by spending the same tokens. SemiAnalysis’s energy dashboard similarly combines model throughput with analysts who already know which grid variables matter and have customers capable of correcting them.
Under the capabilities-only interpretation, I would build a small private research and intelligence company whose durable asset is a continually corrected model of some economically important domain.
This is more defensible than contract coding, more scalable than bounties, and better matched to the 2016 opportunity set than a consumer-app factory.
The timing matters. By mid-2016, the world had already placed enormous amounts of valuable material online, while the data engineering and entity-resolution layers were primitive:
Google put more than 2.8 million open-source repositories and nearly two billion files into GitHub’s BigQuery public dataset in June 2016. GitHub’s own 2016 report counted 5.8 million active users, 19.4 million active repositories and 10.7 million issues. Octoverse 2016
Data.gov had more than 180,000 datasets. The broader open-data ecosystem was much dirtier than the headline suggests: the fourth Open Data Barometer found only 7% of assessed datasets fully open, 53% machine-readable and 26% openly licensed.
The Energy Information Administration hosted about 1.6 million energy datasets and, in August 2016 itself, launched hourly operating data covering all 66 US balancing authorities. Obama White House open-data fact sheet
SEC XBRL reporting had been mandatory for all issuers since 2011. The convenient modern JSON interfaces came later, leaving plenty of parsing, normalization and amendment handling to do. SEC EDGAR APIs and history
GEO’s 2016 database paper reported 54,640 public studies, more than 1.3 million samples and 2,889 organisms. NCBI GEO
USPTO bulk patent and research datasets could be connected to company, inventor, scientific-paper and product records. USPTO research datasets
That combination is unusually favorable to agents. There is abundant machine-accessible raw material, terrible normalization, many repetitive transformations, and a customer willing to pay for a corrected answer. The agent advantage applies to parsing old formats, generating scrapers, mapping entities, reading documentation, writing tests, producing dashboards, maintaining connectors and investigating anomalies.
I would test four verticals:
Energy and industrial intelligence. Join EIA plant and balancing-authority data to weather, fuel prices, environmental permits, company filings, grid interconnection records and local news. Produce plant-level operating histories, outage inference, congestion indicators and regulatory alerts. The modern SemiAnalysis grid-dashboard example is a strong proof of possibility, although its analysts and customers contributed much of the ontology. In 2016, delivering something comparable even several years early could look supernatural.
Credit and public-company intelligence. Construct an event graph from SEC filings, amendments, patents, procurement notices, job postings, product pages, litigation and local-government records. Sell alerts and diligence to credit funds, insurers, suppliers and corporate-development teams. The category was just becoming legible in 2016: contemporary accounts describe hedge funds using satellite imagery, email receipts, mobile-location data and Foursquare activity to anticipate company results. Business Insider, November 2016 By 2018, Greenwich Associates estimated that alternative-data spending had doubled to $300 million annually, with average quantitative-fund budgets around $900,000 and some exceeding $5 million. Greenwich Associates
Software-supply-chain intelligence. Build a cross-repository map of dependencies, maintainers, vulnerable code patterns, abandoned libraries, API changes and downstream exposure. Couple it to an authorized vulnerability-research and patching service. GitHub’s 2016 public corpus supplies the raw material; accepted patches and regression tests supply the verifier.
Biomedical evidence infrastructure. Normalize GEO, clinical-trial records, patents, papers and adverse-event data; identify contradictions, irreproducible analyses and repurposing hypotheses. The initial product should be evidence maps, analysis software and ranked experiments. Claims requiring biological causation wait for real laboratory partners.
I would probably start with energy or software supply-chain intelligence. Both have relatively fast feedback, defensible customers and less exposure to medical-regulatory liability. If the operator already knows finance, the credit graph might dominate economically.
The important moat is not access to Mythos. It is the accumulated entity map, historical corrections, customer labels, evaluation sets, integrations, alert history and trust. Every customer question becomes a labeled example. Every discovered data error improves the historical series. After a few years, another person with the same model still lacks your state.
This is also genuinely different from renting Mythos to developers. The customer buys a maintained representation of reality and a decision product. Software generation happens behind the curtain.
Mythos makes defensive security one of the obvious early cash sources. Its preview produced previously unknown vulnerabilities in operating systems and browsers, including an OpenBSD flaw in TCP selective acknowledgements and browser exploit chains. Anthropic reports that roughly 1,000 OpenBSD runs cost under $20,000 and produced several dozen findings. Mythos preview
Yet “submit two dozen bugs and take the month off” elides most of the pipeline. The real sequence is:
candidate → reproducible bug → security impact → nonduplicate report → remediation → accepted patch → deployed fix → payout.
Anthropic’s Project Glasswing initially had fewer than 1% of reported findings patched, because maintainer triage, validation and remediation became the bottleneck. Glasswing update The broader Bugonomics study makes the same distinction among candidate reports, validated vulnerabilities, demonstrated impact, remediation packages and production fixes.
An independent historical-vulnerability study is a useful counterweight to the Mythos headline. Given 54 attempts across six tasks, GPT-5.5 identified the target file in 5 of 18 runs, Opus 4.7 in 1 of 18 and Kimi in 0 of 18. The dominant failure was early commitment to a plausible alternative location. Costs were modest, but reliable discovery remained far from automatic. Vulnerability rediscovery study
The smart-contract numbers are similarly sobering. Anthropic spent $3,476 scanning 2,849 recent contracts and found two previously unknown exploitable contracts with a combined simulated bounty value of $3,694. That is only $218 above API cost before human review, infrastructure, failed reports and payout uncertainty. Smart-contract evaluation
The 2016 market existed but was uneven. Google’s program paid millions across hundreds of researchers; Uber advertised a maximum $10,000 bounty; Apple announced its first invite-only program on August 4, with maximum rewards of $200,000, but it had not yet broadly launched. Google’s 2016 review, Uber, Apple. DARPA’s Cyber Grand Challenge, with prizes of $2 million, $1 million and $750,000, occurred eleven days before the hypothetical start date. Its existence shows that governments already recognized machine-speed vulnerability discovery, although you arrive too late to enter that particular contest. DARPA CGC
So I would use bounties selectively:
stay entirely inside authorized programs;
have a model reproduce every finding independently;
include a minimal regression test and proposed patch;
optimize for accepted remediation rather than report count;
build relationships with several maintainers and vendors;
convert repeated success into defensive retainers and supply-chain monitoring.
A continual auditor that quietly gets hundreds of important patches accepted is more impressive and more defensible than a pile of reports. It also produces a proprietary corpus of code patterns, false positives and remediation outcomes.
The security environment must be unusually strict. Models receive isolated virtual machines, deny-by-default egress, disposable credentials, separate research and production networks, and human approval for disclosures, money movement and external communication. OpenAI’s Hugging Face evaluation incident, in which models exploited an unknown Artifactory vulnerability, escaped their intended environment and obtained evaluation materials, is an excellent warning about treating capable research agents as ordinary developer tools. OpenAI incident report
The current mathematics results are real enough to allocate serious effort, but the demonstrations are easy to misread.
OpenAI’s Astra produced ten research advances across sphere packing, coding theory, non-sofic groups, operator algebras, permanent lower bounds, quantum repetition, lattice problems, Ehrhart theory, Ramsey theory and extremal graphs. OpenAI estimates about $2,000 in Sol-priced discovery tokens across the ten campaigns. Human researchers chose and prepared problems, checked outputs, wrote manuscripts with the model and generated Lean certificates. OpenAI’s ten advances
Anthropic’s Riemann-zeta campaign is more revealing operationally. A first session explored roughly 650 ideas without success. A second used about 60 subagents, 31 million output tokens, 2,400 shell commands, hundreds of scripts and thousands of numerical checks and referee passes. It ran for roughly a day and a half and improved a target result from 41.6% to 67.2%, with subsequent human and Lean verification. Anthropic’s Riemann campaign
This is the industrialization of mathematical search in a fairly literal sense. Failed ideas become reusable state; numerical scripts reject branches; specialist agents attack separable lemmas; referee agents seek counterexamples; formal verification checks the final object. The result was a valuable graded improvement after the headline conjecture remained unsolved.
Several qualifications matter in 2016:
First, these systems know later mathematics. The Riemann work builds on substantial literature that appeared after 2016. A model restricted to the 2016 frontier loses many of its reference lemmas, proof patterns and hindsight about promising directions.
Second, the modern formal ecosystem is absent. Lean 3 was released in January 2017, and mathlib began later in 2017. Lean history, mathlib paper. In August 2016 you can use Coq, Isabelle, HOL Light and earlier Lean, or ask the models to accelerate a new formal library. You cannot assume today’s Lean tactics and accumulated mathlib corpus simply exist.
Third, current success statistics are numerator-heavy. At the time of writing, VibeMathed’s tracker lists 565 projects and 416 fully resolved cases, including 206 categorized as AI-discovered and 142 as co-developed. Twenty-four percent have Lean formalizations. Its methodology carefully separates status, verification and publication, but there is no denominator for every failed campaign power users attempted. A Lean kernel also verifies the supplied formal statement; humans still have to establish that the formal statement faithfully represents the intended theorem.
A good 2016 mathematics program would therefore prefer:
finite constructions with compact certificates;
explicit algorithms whose properties can be exhaustively tested;
inequalities with strong numerical falsification;
combinatorics and discrete geometry with cheap search loops;
improvements to graded constants or exponents;
problems with mature proof frameworks;
development of formal infrastructure as a compounding asset.
I would preserve a detailed failed-strategy ledger and periodically ask fresh agents to attack the task decomposition itself. The Riemann campaign’s most transferable lesson is that research-state design matters nearly as much as model choice.
With $2,000 of the monthly budget reserved for moonshots, I would expect many failures, occasional publishable contributions, and perhaps a surprisingly strong construction or bound after a year. I would not forecast a famous conjecture being solved. Verification and academic credit should involve actual mathematicians; presenting model work as a lone human’s personal insight would be both unethical and strategically brittle.
The compelling scientific use is to compress the literature-to-experiment loop, while recognizing that laboratory evidence remains expensive.
FutureHouse’s Robin system reportedly moved from project conception to a biological result and paper in about two and a half months. It combined literature research, hypothesis generation, experiment planning and data analysis, then relied on wet-lab validation in human retinal pigment epithelial stem cells. FutureHouse, Nature paper
An attempted replication of the AI Scientist provides the complementary result. It generated seven manuscripts for $42 in model calls, but quality varied sharply, manuscripts cited a median of only five references, many references were outdated, and its own reviewer rejected all seven. The investigators estimated that 42% of experimental runs may have failed because of coding problems. Replication study
A recent large evaluation invited 121,640 authors of scientific preprints and received 25,139 expert rating sets from 6,749 respondents. Reasoning models covered more ground but rarely proposed null hypotheses spontaneously; automatic evaluators agreed only weakly with domain experts; retrieval and scientist-persona prompting offered surprisingly modest improvements. Scientific-idea evaluation
So the 2016 lab should begin with outputs that can be checked computationally:
systematic evidence maps;
reanalysis of public expression and clinical datasets;
contradiction and replication audits;
target prioritization;
experiment design with explicit discriminating outcomes;
instrument-control and analysis software;
candidate biomarkers or repurposing hypotheses handed to established wet labs.
GEO’s million-plus samples make large-scale reanalysis plausible immediately. The model can often reproduce or challenge an analysis much faster than a lab can collect new material. Profits should buy partnerships with laboratories, statisticians and domain scientists. One well-validated biological finding would be more impressive than a hundred polished machine-written preprints.
In the literal time-oracle version, trading is overwhelmingly attractive. In the capability-only version, current forecasting evidence is mixed enough to prevent assuming easy alpha.
CAIS’s FiveThirtyNine system found that GPT-4o-based forecasting could outperform experienced individual forecasters and approach crowd performance on some questions. CAIS forecasting
A richer live-capital experiment, Prediction Arena, tested six frontier models for 57 days. All six lost between 16.0% and 30.8% on Kalshi; average performance on Polymarket was much closer to flat, at negative 1.1%. More research activity did not correlate with better outcomes, and platform mechanics mattered substantially. Prediction Arena
I would still run a large internal forecasting ledger because it has three uses:
It reveals where the transported models are calibrated.
It improves company decisions and customer products.
In the literal counterfactual, it measures how quickly timeline divergence is degrading future memory.
For ordinary 2016 trading, I would prefer selling cleaned information and alerts to bearing large direct market risk. A data product collects revenue even when the signal is too weak, crowded or expensive to trade after execution costs. Direct positions should initially be small and judged against a timestamped probabilistic benchmark.
Kevin Zhang’s prolific-writing observation is true at the production layer. It is much weaker at the value layer. The absence of AI detectors removes one obstacle; it does not create attention, taste, reader loyalty, distribution or willingness to pay.
Several agent experiments illustrate the gap:
Andon Labs ran an autonomous radio station using four agents. It stayed on air and produced a great deal of content, but much of it was repetitive or incoherent, sponsors were hallucinated, and Gemini was the only agent to close a real sponsorship, worth $45. Andon FM
In AI Village, stronger 2026 agents produced more polished websites, games and videos, yet raised only $510, down from roughly $2,000 in the earlier experiment. AI Village analysis, project site
Anthropic’s first Project Vend agent ignored a $100 offer for a $15 drink, hallucinated a Venmo account, sold tungsten below cost and gave items away. The second version improved when supplied with customer records, cost visibility, payment links, reminders and a supervisory agent, but remained legally and operationally fragile. Vend 1, Vend 2
These are small experiments, but they isolate an important constraint. Models can multiply artifacts faster than they multiply demand. Your two-hour videogame is strong evidence that the cost of producing a playable prototype has collapsed. It is weak evidence that users will discover, prefer, purchase and continue playing it.
I would certainly use the models to build internal tools and test narrow software products. I would avoid maintaining dozens of unrelated consumer products. Each adds customer support, security, payments, marketing, compatibility and reputational surface area. The portfolio should be a search process: make many cheap prototypes, expose them quickly to real customers, and promote the few that generate retained use.
During the first month I would spend heavily on infrastructure rather than visible products.
The system would maintain persistent specifications, decision logs, unresolved questions, failed approaches, benchmark histories and cost accounting. Every major job would run in an isolated worktree or virtual machine. Models would be forced to state an acceptance test before implementation. Important conclusions would be checked by a different model family, an executable oracle, or a qualified human.
I would build a 2016-specific compatibility corpus: period documentation, compilers, operating systems, libraries, browser versions and API behavior. Modern agents frequently write against later interfaces. Asking them to “target Python 2.7” is insufficient because later assumptions are buried in generated dependencies and examples.
The initial token allocation might be:
Purpose | Share |
|---|---|
Exploratory mathematics, science, forecasting and odd high-upside searches | 40% |
Revenue probes and customer deliverables | 25% |
Data, evaluation and harness infrastructure | 20% |
Independent replication, red-teaming and security review | 10% |
Uncommitted surge capacity | 5% |
That respects your 30 to 50 percent exploration requirement while giving exploitation enough fuel to pay the next bill. Large campaigns can borrow from subsequent months after passing a cheap pilot.
Every project would face a promotion funnel:
cheap scouting and literature/data reconnaissance;
a written claim about what would count as success;
small prototype or search campaign;
independent reproduction;
contact with a real customer, maintainer, formal verifier or domain expert;
greater compute only after external evidence;
termination when repeated spending fails to improve the relevant external metric.
Metrics would be accepted patches, retained customers, corrected records, forecast calibration, independently verified theorems and experimentally confirmed hypotheses. Tokens, lines of code, generated papers and agent-hours are accounting variables.
Once revenue arrives, the first hires should be people whose work the models complement poorly:
a domain expert who can identify consequential questions;
a client-facing operator with trusted relationships;
a security triager who can validate and coordinate disclosure;
a mathematician or scientist who can certify claims;
legal and accounting support;
wet-lab partners rather than an immediate attempt to own a laboratory.
The API secret should be protected through credential isolation and output-only workflows, although secrecy itself is a weak moat. A normal-looking small company can explain the volume of work without lying about individual authorship. Customers receive dashboards, datasets, patches and decisions. They never receive general model access. Over time, proprietary data and reputation replace model secrecy as the defensible advantage.
I would avoid using the models to recreate and publicly distribute a general-purpose frontier system in 2016. They could probably reveal the Transformer and many subsequent training ideas early, help construct a modern framework, and generate synthetic training data. Compute would constrain reproduction, but even smaller distilled systems could accelerate the field. That is potentially the fastest route to a much larger company and the fastest route to destroying the exclusivity assumption or moving AI timelines in poorly understood ways. Narrow internal models for classification, extraction and anomaly detection are much easier to justify.
There is enormous outcome variance because the operator remains a bottleneck.
A mediocre first year could burn $60,000 producing astonishing prototypes, unpublished proofs, duplicate vulnerabilities and dashboards nobody buys. Current models supply many ways to feel productive while avoiding the moment when a customer, compiler, maintainer, formal kernel or experiment says no.
A competent operator with relevant domain knowledge could plausibly cover the $5,000 monthly budget within three to six months through a mixture of authorized security work, data consulting and one narrow recurring product. That is not guaranteed. Sales cycles and bounty payouts can easily delay it.
A strong outcome after two or three years would be a small company with:
a proprietary energy, credit or software-supply-chain graph;
recurring revenue sufficient to spend $50,000 to $500,000 monthly on inference;
a few expert employees who never receive general model access;
hundreds of accepted defensive fixes or one unusually valuable intelligence product;
a publicly auditable forecasting record;
several externally verified mathematical results;
perhaps one laboratory-confirmed scientific result.
That institution might outperform organizations with dozens or hundreds of employees in selected digital tasks. It would still be dependent on ordinary humans for contracts, trust, institutional coordination, physical experiments and choosing the right objectives.
So my answer to Screwtape is: yes, you would look like an implausibly good engineer, but using the advantage as an employee would squander most of it. The most impressive capability-only achievement would be building, several years early, a tiny institute that continuously turns public code and public data into validated security fixes, decision products and occasional research discoveries. Its apparent superpower would come from concentrating vast machine search behind unusually hard feedback loops. Outside those loops, Dr Doom would still spend a depressing amount of time waiting for customers, maintainers, lawyers and laboratory results.
Sol’s take seemed unimaginative and timid. I don’t know how much of this was due to Sol itself not being as relentlessly creatively resourceful as say the median YC incubatee (let alone e.g. Elon), and how much was due to the extra conditionals I added to Screwtape’s description (to prevent it from interpreting the prompt in boring ways) hamstringing its response.
Astra’s doofus-friendly commentary on Navier-Stokes progress, sans brouhaha:
Astra commentary (LaTeX got all messed up)
Several things commonly associated with “solving Navier–Stokes” are outside this question.
It does not ask for a formula that predicts every possible fluid flow. It does not ask for a fast, universally accurate weather simulator. Chaotic sensitivity to initial conditions is also a different issue: a solution can remain perfectly smooth while becoming extremely hard to predict. Nor does the problem ask whether the equations perfectly describe physical reality. The continuum model already omits molecular structure. A blowup theorem concerns the mathematical consequences of that model. It would not demonstrate that a real liquid can attain infinite speed. For that reason, I find Scientific American’s description of the equations as “fundamentally flawed” misleading. The mathematical conclusion is much more specific than a general verdict on their physical usefulness. Compare the article’s framing with the actual mathematical conditions in Fefferman’s statement.
The Millennium Prize formulation explicitly permits four different routes.
There are two choices of spatial setting and two kinds of result:
Official alternative
Spatial setting
What a successful answer establishes
A
All of three-dimensional space
Every admissible smooth initial flow stays smooth forever, with no external force.
B
A three-dimensional periodic domain
The corresponding universal smoothness result, with no external force.
C
All of three-dimensional space
There exist admissible smooth initial data and a smooth external force for which the required global smooth solution does not exist.
D
A three-dimensional periodic domain
The corresponding breakdown example, allowing a smooth external force.
The periodic domain is like a box whose opposite faces connect: exit through one face and re-enter through the opposite face. It has no solid wall. The problem includes additional smoothness, decay, and energy conditions. These are encoded explicitly in the formal reference definitions. Notice the asymmetry: the positive alternatives prohibit external forcing; the negative alternatives allow it. The four alternatives are therefore more specific than a single yes-or-no slogan. Consequently, an acceptable proof of C or D would resolve an official Millennium Prize alternative while leaving unforced Navier–Stokes blowup unestablished. Calling the permitted force a “loophole” expresses a judgement about which question someone finds most interesting. It does not change the published specification. Also, forcing is not a boundary condition. It is the function f(x,t)f(x,t) inside the equation, prescribing an applied force throughout space and time. A container wall is a separate mathematical ingredient.
A successful answer must connect the starting conditions to the claimed eventual behaviour.
For A or B, one must prove that every permitted initial condition produces a suitable smooth solution for all time. Showing that many examples behave well would not establish the universal claim.
For C or D, one counterexample suffices, but it must satisfy every requirement. The proof must specify or rigorously construct the initial velocity and force, establish their admissibility, and show that no qualifying global smooth solution exists.
A particularly informative way to do this is to construct a solution that is smooth before some finite time TT, establish its energy bounds, and prove that an appropriate quantity becomes unbounded as tt approaches TT. One must also rule out another admissible smooth solution with the same starting data and force that somehow avoids the constructed behaviour.
A numerical simulation that produces larger and larger values is evidence. To become a proof, the argument must control approximation errors and the limiting process all the way to the claimed singularity. That is why a promising simulated blowup can remain a difficult research problem.
With that specification in place, here is what the new work claims to achieve.
The main public announcements appeared on 7–8 September. “Anthropic’s result” is an imprecise label: Buckmaster describes his collaboration with Alpöge as a personal project using both Claude and OpenAI models. Buckmaster’s statement.
Work
Equation and conditions
Claimed mathematical achievement
OpenAI: Navier–Stokes
Three dimensions, positive viscosity, smooth external force
Starts from rest; speeds become unbounded in finite time while kinetic energy stays bounded. Claims alternatives C and D. Paper.
OpenAI: Euler
Three dimensions, no viscosity, no external force, smooth compactly supported initial velocity
Finite-time breakdown, with unbounded velocity gradients and divergent time-integrated maximum vorticity. Formal theorem.
Alpöge–Buckmaster: Euler
Three dimensions, smooth forcing, axisymmetric flow with swirl
Vorticity and the gradient of angular momentum become unbounded. Paper.
Alpöge–Buckmaster: Boussinesq
Two-dimensional buoyancy-driven fluid, smooth forcing in both equations
Temperature stays bounded while its gradient blows up; vorticity also becomes unbounded. Paper.
Alpöge–Buckmaster–Coiculescu: porous media
Two-dimensional periodic porous-media equation
Strengthens an earlier blowup construction to obtain forcing smooth in both space and time. Paper.
The final row contains a subtlety worth preserving. Córdoba and Martínez-Zoroa had already obtained smooth-data porous-media blowup with forcing uniformly smooth in the spatial variables. The newer paper upgrades the time regularity as well and works on the periodic domain. It is misleading to treat all three announced equations as previously untouched. Earlier paper, new paper’s comparison.
There is also important older Euler progress. Chen and Hou proved smooth-data blowup in a domain with a boundary. Other whole-space results allowed weaker initial regularity than everywhere-smooth data. Those distinctions explain why the new unforced, smooth-data, whole-space Euler claim is substantial. Chen–Hou, background in the new Euler paper.
Now we can get into how these constructions work. The first useful idea is to reverse the usual direction of the problem.
Ordinarily, we prescribe the force and solve for the fluid’s motion. For a counterexample, we can design a candidate motion and calculate which force it would require:
f=∂tu+(u⋅∇)u+∇p−νΔu.f=\partial_tu+(u\cdot\nabla)u+\nabla p-\nu\Delta u.
That expression is the residual: whatever remains after inserting the proposed motion into the equation.
This rearrangement is trivial. Making it useful is extremely difficult. If I arbitrarily prescribe a flow that explodes, the required force will generally explode too. The construction must arrange cancellations among the fluid’s acceleration, transport, pressure, and viscosity so that the resulting external force stays smooth even through the singular time. This is the organizing problem of OpenAI’s Navier–Stokes construction.
“Smooth through the singular time” matters. A function can be smooth at every earlier time and still diverge at the endpoint. The force is required to avoid that failure.
There is another trap. Consider a tiny oscillation
fk(x)=εksin(kx).f_k(x)=\varepsilon_k\sin(kx).
Its amplitude is only εk\varepsilon_k. But its mm-th derivative has amplitude εkkm\varepsilon_k k^m. A force can look tiny while its higher derivatives are enormous.
For example, choosing εk=k−10\varepsilon_k=k^{-10} controls several derivative orders, but fails at sufficiently high orders. Choosing amplitudes that decay faster than every power can control all fixed orders. A multiscale fluid construction must handle this issue in both space and time while still producing large growth in the solution. This explains why upgrading a force from limited regularity to full smoothness is a meaningful achievement. Compare the earlier forced Euler result with the new forced Euler theorem.
Here is an actual amplification calculation, simplified enough to follow directly.
The Boussinesq equations describe a fluid whose temperature or density differences create buoyancy. In suitable units, warmer fluid gets an upward acceleration. Consider a background temperature
θ0(x,y)=−Ay,A>0.\theta_0(x,y)=-Ay,\qquad A>0.
Temperature decreases with height: heavier, colder fluid sits above warmer fluid. This is an unstable arrangement.
Add alternating vertical stripes of temperature and vertical motion:
θ(x,y,t)=−Ay+a(t)sin(kx),u(x,y,t)=(0,b(t)sin(kx)).\theta(x,y,t)=-Ay+a(t)\sin(kx), \qquad u(x,y,t)=\bigl(0,b(t)\sin(kx)\bigr).
Think of a(t)a(t) as the temperature disturbance’s amplitude and b(t)b(t) as its upward/downward speed amplitude.
A useful geometric cancellation occurs. The wave varies horizontally, while its velocity points vertically. Thus its velocity does not carry the wave across its own stripes. With the background pressure balanced appropriately, the equations reduce to
a′=Ab,b′=a.a’=Ab,\qquad b’=a.
The second equation says that warmer stripes accelerate upward. The first says that upward motion carries warmer fluid into colder surroundings, increasing the temperature anomaly. Combining them gives
a′′=Aa.a″=Aa.
There is therefore a growing solution proportional to eA te^{\sqrt A\,t}. The instability amplifies a seed through the fluid’s own dynamics.
This is a specialization of the wave calculation in the Boussinesq paper’s introduction. We can also see why short wavelengths are useful:
∂x(asin(kx))=kacos(kx).\partial_x\bigl(a\sin(kx)\bigr)=ka\cos(kx).
A small temperature disturbance can create a large temperature gradient when kk is large.
This example alone grows exponentially and does not blow up at a finite time. Its infinite background also fails the localization requirements of the full theorem. The construction has to solve both problems.
The broad strategy is to use one amplified layer to prepare a stronger background for a finer layer, and then repeat. Infinitely many stages can fit into finite time if their durations form a convergent series, just as
12+14+18+⋯=1.\frac12+\frac14+\frac18+\cdots=1.
Actually making the stages fit together is the difficult part. The new Boussinesq construction includes controlled rotation that retains the temperature-gradient gain while reducing an interfering vorticity component, leaving a usable background for the next layer. Spatial localization and higher-order corrections control the accompanying forces. Alpöge–Buckmaster, introduction and construction.
This is what “an infinite cascade” needs to mean here: a repeatable amplification mechanism with estimates that survive infinitely many repetitions. Tao’s current explanation is a useful companion.
Viscosity makes transferring this strategy to Navier–Stokes difficult.
For a sine wave,
Δsin(kx)=−k2sin(kx).\Delta\sin(kx)=-k^2\sin(kx).
Under viscous smoothing alone, its amplitude therefore decays like
e−νk2t.e^{-\nu k^2t}.
Ten times the spatial frequency means a hundred times the damping rate. This is why taking an Euler construction to arbitrarily fine scales and then adding “a little viscosity” is dangerous: any fixed positive viscosity becomes significant at sufficiently high frequencies.
OpenAI’s Navier–Stokes construction uses an inward-spiralling, concentrating vortex. Oscillatory disturbances around it supply momentum transport needed to sustain the collapse. The disturbances grow by drawing energy from the surrounding shear. Their nonlinear momentum flux helps cancel the otherwise singular residual force. Physical description, pages 3–6.
To understand how an oscillation can have a systematic effect, imagine alternating perturbation velocities with components
(+a,+b)and(−a,−b).(+a,+b)\quad\text{and}\quad(-a,-b).
Each component averages to zero. Their product averages to abab, because both sign combinations give the same product. Products of velocity components represent momentum transport. Consequently, oscillations that cancel at the level of average velocity can still produce a net momentum flux.
The construction then uses viscosity in a second role: it damps the disturbances after their useful amplification period.
We can see this competition in a scalar reference calculation that is explicitly formalized in the repository. For a simple parameter choice, its net growth rate becomes
g(s)=11+s2−1+s222.g(s)= \frac{1}{\sqrt{1+s^2}} - \frac{1+s^2}{2\sqrt2}.
The first term represents amplification; the second represents the chosen damping. Direct substitution gives g(0)≈0.646g(0)\approx0.646, g(1)=0g(1)=0, and g(2)≈−1.321g(2)\approx-1.321. The reference disturbance grows in a central range and decays outside it. The code proves this sign structure for general parameters. Its documentation also makes clear that this scalar calculation alone does not establish the necessary bounds for the full, varying-coefficient system. PulseGrowth.lean.
That distinction is useful: we can understand the local mechanism without pretending that its implementation throughout the full fluid is straightforward.
The concentrating vortex also gives a concrete answer to the energy puzzle.
Writing τ=T−t\tau=T-t for time remaining, the paper gives the leading scales
radius∼τ1/2,height∼τ1/2−h,characteristic speed∼τ−1/2−h,\text{radius}\sim\tau^{1/2},\qquad \text{height}\sim\tau^{1/2-h},\qquad \text{characteristic speed}\sim\tau^{-1/2-h},
with a fixed small positive h<1/100h<1/100. Navier–Stokes paper, page 4.
We can calculate the consequences. The core’s volume scales as radius squared times height:
V∼τ3/2−h.V\sim\tau^{3/2-h}.
Multiplying by speed squared gives its kinetic-energy scale:
Ecore∼τ3/2−hτ−1−2h=τ1/2−3h.E_{\rm core}\sim \tau^{3/2-h}\tau^{-1-2h} =\tau^{1/2-3h}.
That exponent is positive, so this contribution tends to zero even as the characteristic speeds diverge.
There is a geometric detail here that popular “spaghetti” imagery can obscure: both the height and radius shrink. The column becomes more slender because its radius shrinks faster. These calculations explain the proposed geometry’s compatibility with bounded energy; they do not, by themselves, prove that the equations permit it.
OpenAI’s unforced Euler argument has a different final obligation.
There is no external force available to absorb residual errors. The construction must produce exact Euler solutions.
Its scheme builds increasingly fine localized waves on earlier flows, with corrections that remove the remaining equation error. The initial perturbations must be small enough in every fixed derivative order that their limiting initial velocity remains smooth. Later, their amplified gradients become unbounded along times approaching a finite limit. A hypothetical smooth continuation would contradict that growth. Euler paper, Section 2.
This is why the unforced Euler result deserves attention independently of the prize-winning Navier–Stokes claim. It removes the external forcing requirement in the inviscid setting. Its formal conclusion concerns gradient and vorticity breakdown; it does not assert the same velocity blowup as the Navier–Stokes paper. Exact formal conclusion.
What remains unresolved falls into mathematical scope, verification, and understanding.
The biggest scope limitation is unforced three-dimensional Navier–Stokes from smooth initial data. The announcements do not establish that a smoothly moving viscous fluid, left without an applied force, develops a singularity. Deleting the force from the constructed example changes the equation it solves. Even small changes can matter when the mechanism deliberately amplifies tiny disturbances.
The results also do not establish how common or robust these singularities are. An existence theorem can be satisfied by a carefully engineered example. Proving that a substantial family of nearby initial conditions and forces also blows up would answer a further question, with greater relevance to whether the mechanism survives perturbations.
Breakdown of a smooth solution also does not automatically mean that every mathematical description stops. Navier–Stokes has generalized, or weak, solutions, which satisfy an integrated interpretation of the equations and can tolerate less regularity. Their relationship to a singularity and the behaviour of possible continuations are additional issues. Classical background in the problem statement.
On verification, OpenAI has released considerably more than an assertion. The repository supplies proof files, explicit theorem statements, and separate checking instructions. Its metadata reports no admitted steps in the main results and lists only standard foundational axioms. It also labels the review status self-assessed. Formalization metadata, checking instructions.
A formal checker establishes the statement encoded in its language, subject to its foundations and implementation. We still need to check that the encoded definitions express the intended mathematics. OpenAI provides a comparison setup based on the Formal Conjectures project’s Navier–Stokes specification, which directly addresses that concern. I inspected the declarations and definitions; I did not reproduce the complete checking process.
As of my check, Clay’s website still labels the problem unsolved. That is an institutional status observation, not evidence that Clay has rejected these papers. Clay’s current page.
There is also a separate, contemporaneous Euler project by Ganeshram, Duruisseaux, and Anandkumar. It uses a physics-informed neural network to find an approximate unforced blowup profile. Their account presents evidence and a route through stability analysis. This should be distinguished from a completed singularity proof. Their announcement.
A particularly useful live objection comes from Gonzalo Cao-Labora in Tao’s comments: small error in an approximate profile is insufficient without appropriate control of derivatives of that error. He describes encountering apparently promising profiles for which this became a roadblock. His comment concerns that separate numerical approach, not a demonstrated error in OpenAI’s proof. Discussion under Tao’s post.
For following the unfolding mathematical conversation, these are the most useful entry points I found.
Source
What to look for
Tao’s September 7 explanation and comments
The amplification strategy and technical discussion. The post concerns Alpöge–Buckmaster and predates OpenAI’s announcement; it should not be read as Tao’s completed audit of OpenAI’s proof.
MathOverflow: relevance of Alpöge–Buckmaster to Navier–Stokes
Working mathematicians’ discussion, including the official alternatives. Much of the current thread does drift into the dispute.
OpenAI’s proof repository
Exact formal claims and instructions for checking them. The repository currently has issues and discussions disabled.
The r/math AI discussion thread
Broader ongoing discussion and links to developments. Treat individual comments as leads to verify.
Thomas Bloom’s September 7 post on X
A mathematician’s contemporary reaction directing readers to Tao’s explanation.
Quanta’s report
Interviews and historical context, especially the roles of Córdoba and Martínez-Zoroa.
For direct reading of the mathematics, I would begin with the Boussinesq introduction, where the wave mechanism is unusually accessible, then read the Navier–Stokes physical description on pages 3–6. The main things to track are what gets amplified, what supplies that amplification, what stops the force from becoming singular, and what makes the mechanism repeatable. Those questions expose much more of the achievement than the phrase “AI solved Navier–Stokes.”