See my other comment in this thread for actual AI alignment thoughts, but as a former aerospace engineer myself (albeit not a very good one), I thought it would be fun to speculate on “Would such a cheerful innocent ever succeed at landing a space probe? It seems crazy to assign them a real-world success probability as high as 10%.”
In the very early years of cubesats (very small satellites built from off-the-shelf components, sometimes as university projects), through around 2009, about half of all cubesats launched into space were “dead on arrival”, ie no communication was ever made with them after launch, or suffered “infant mortality” (communication was lost within days of launch). Here is a blog post with lots more detail on beginner cubesat failure rates, causes, etc (also featuring a truly unexpected Harry-Potter-and-the-methods-of-aerospace-engineering theme throughout the later section headings).
In later years, this number appears to have improved (from 50% to around 20%, which is still crazy high), but I think this seeming improvement is mostly due to a combination of: 1. a few serious companies, like Planet Labs, launching large numbers of duplicate cubesats that they worked hard to get right, and 2. universities / tiny companies / etc being able to buy increasingly complete “off the shelf” cubesats based on components that had increasingly strong track records of prior flights, which are effectively retries (not first-critical-tries) on behalf of the company making those components.
If you subtract out the serious companies full of serious aerospace engineers, and the effective retries, the failure rate of the remaining “truly naive attempts” from people who eg have barely even read blog posts warning of potential dangers like the one I linked earlier, is definitely way over 50%, maybe 80%… obviously the cutoff of what you count as a truly naive attempt is subjective; at the limit you are just filtering for “the dumbest most unprepared cubesat teams ever” which surely have a failure rate of 100%.
But Eliezer’s analogy wasn’t positing a team of the dumbest, most unprepared people ever. He was positing a team of smart, well-resourced people who are in a certain sense trying hard, but nevertheless also posess suicidal naivete about the dangers of space probe design. What success probability would such a team have of launching a working cubesat on the first try?? idk, 10% doesn’t seem crazy; even people who are suicidally naive (ie, totally failing to consider failure cases and recovery modes and unknown-unknowns and being paranoid, but otherwise doing good-quality engineering if such a thing is even philosophically concievable: doing some customary tests of their satellite on the ground, etc, just totally failing to think for themselves about how things could actually go wrong) would probably luck into creating a working cubesat (that isn’t just a carbon copy of some earlier project that worked) like 20% − 40% of the time.
BUT, Eliezer didn’t say “make a cubesat”, lol. That’s like the easist possible space task!! Anyone can make Sputnik 1; the hard part is obviously making the rocket… and then you have to “succeed at landing a space probe”, presumably on Mars or the moon. Yeah this is starting to look completely impossible.
Getting to test the rocket with unlimited retries in the atmosphere actually plausibly gets you most of the way to a working orbital rocket—as far as I’m aware Starship has only done suborbital flights so far, and it’s shaping up as a pretty serious, mostly-finished rocket (albeit these flights have gone into space, but they could’ve done similarish flights that technically stayed in the atmosphere if they had to). In real life of course nobody gets unlimited retries, see herefor the assorted failure modes that doomed all the different attempted flights of the Soviet N-1 moon rocket, which flew four times and blew up four times—featuring phrases like:
“One unforeseen flaw was that [the rocket’s command computer] operating frequency, 1000 Hz, happened to perfectly coincide with vibration generated by the propulsion system, and the commanded shutdown of Engine #12 at liftoff was believed to have been caused by pyrotechnic devices opening a valve, which produced a high-frequency oscillation...”
“The engine control system would also be reworked, increasing the number of sensors from 700 to 13,000.”
“One of the largest accidental artificial non-nuclear explosions in history.”
But even if you test everything you can in the atmosphere, your rocket probably still just immediately fails on some aspect of its uppermost stage that’s supposed to push your space probe to the moon/mars. Upper stages are way harder than cubesats: you have to deal with propulsion systems (with valves that can freeze in weird ways in space, and you can’t really “test a propulsion system in vacuum” like you can put a satellite in a vacuum chamber), you have to actually orient and point the proper direction (cubesats can just tumble), you have moving parts like fairings and decouplers that again might possibly behave weirdly in space and are not trivial to test in fully realistic conditions on the ground, most of the burns have to actually fire at the exact right moment and last for the exact right amount of time otherwise you won’t arrive at the moon/mars (versus if you miss a command on a cubesat because it was resetting or whatever, no biggie, just send the command an hour from now when it’s looped around the earth another time). ChatGPT estimates that in the history of rocketry from the 1980s to now, maybe around 60% of genuinely new upper stages have “basically worked perfectly on the first try”—although obviously those were all built by normal non-naive engineers (indeed, the recent wave of move-fast-and-break-things small-launch startups have a significantly lower hit rate than long-established space programs and traditional defense contractors); maybe fully naive engineers have like a 1⁄5 or 1⁄10 chance of achieving similar outcomes, so maybe 6% − 12%.
Building a moon/mars lander instead of a cubesat is a even more difficult than a rocket second stage, I’d say. Once again you are creating a custom propulsion system, lots of commands have to go off exactly on time (ie during landing), you’ve gotta control your probe’s orientation, etc, but now you have this additional problem of dynamically measuring your distance from unmapped rough ground. Any mistakes in terms of thrust direction / timing now have to be corrected instantly or you hit the ground and die, unlike with an in-space burn where small mistakes can probably be fixed hours later with small correction burns. Also, there’s a good chance your naive-engineer’s plan for dealing with space radiation is basically just “YOLO”, so there’s whatever-percent odds that your ship just dies enroute and whatever half-assed reset procedure exists isn’t enough to get it back. And if you’re landing on mars it’s even worse because you additionally have to worry about heat shields and parachutes and maybe dust messing up your distance measurements, who knows. (If you were a non-naive engineer and knew you had plenty of resources but only got one try, you’d be like “supersonic parachutes are too easy to mess up, we’ll just do heat shield + rockets and it’s fine that the probe will therefore be heavier”, but our naive engineers would miss this.) Similarly a non-naive engineer would probably realize “with infinite resources but only one try we should to to extreme lengths to minimize the number of finnicky moving parts like deployable antennas, solar panels, landing legs, etc, which always fail”, but our naive engineers are just going to have to cross their fingers that their solar panels don’t get stuck in some unexpected way (often it’s hard to perfectly test these sorts of mechanisms because the parts are too fragile to work the same way in earth gravity that they would in zero-g). This is probably 3x harder than making an upper/transfer stage, and for our naive engineers let’s say their odds of success on this task are essentially independent from their odds of success on the upper stage task (since in both cases they’re basically just hoping to luck into avoiding various specific potential failures; the whole concept is they lack the kind of mindset that helps them systematically avoid whole swathes of unknown-unknowns failures), so like 2% − 4%.
So overall I would say maybe 0.3% that a smart and well-resourced but suicidally naive team of engineers lands a space probe on the first try.
The contrast between my gloomy estimate of the success probability for the concrete space-probe thought experiment, versus my relatively optimistic vibe in my on-topic AI alignment comment (tl;dr “come on, what is MIRI’s take on these promising-seeming factors that might help AI go well??”) is left deliberarely unresolved as an exercise for interpretation on behalf of the reader.
See my other comment in this thread for actual AI alignment thoughts, but as a former aerospace engineer myself (albeit not a very good one), I thought it would be fun to speculate on “Would such a cheerful innocent ever succeed at landing a space probe? It seems crazy to assign them a real-world success probability as high as 10%.”
In the very early years of cubesats (very small satellites built from off-the-shelf components, sometimes as university projects), through around 2009, about half of all cubesats launched into space were “dead on arrival”, ie no communication was ever made with them after launch, or suffered “infant mortality” (communication was lost within days of launch). Here is a blog post with lots more detail on beginner cubesat failure rates, causes, etc (also featuring a truly unexpected Harry-Potter-and-the-methods-of-aerospace-engineering theme throughout the later section headings).
In later years, this number appears to have improved (from 50% to around 20%, which is still crazy high), but I think this seeming improvement is mostly due to a combination of: 1. a few serious companies, like Planet Labs, launching large numbers of duplicate cubesats that they worked hard to get right, and 2. universities / tiny companies / etc being able to buy increasingly complete “off the shelf” cubesats based on components that had increasingly strong track records of prior flights, which are effectively retries (not first-critical-tries) on behalf of the company making those components.
If you subtract out the serious companies full of serious aerospace engineers, and the effective retries, the failure rate of the remaining “truly naive attempts” from people who eg have barely even read blog posts warning of potential dangers like the one I linked earlier, is definitely way over 50%, maybe 80%… obviously the cutoff of what you count as a truly naive attempt is subjective; at the limit you are just filtering for “the dumbest most unprepared cubesat teams ever” which surely have a failure rate of 100%.
But Eliezer’s analogy wasn’t positing a team of the dumbest, most unprepared people ever. He was positing a team of smart, well-resourced people who are in a certain sense trying hard, but nevertheless also posess suicidal naivete about the dangers of space probe design. What success probability would such a team have of launching a working cubesat on the first try?? idk, 10% doesn’t seem crazy; even people who are suicidally naive (ie, totally failing to consider failure cases and recovery modes and unknown-unknowns and being paranoid, but otherwise doing good-quality engineering if such a thing is even philosophically concievable: doing some customary tests of their satellite on the ground, etc, just totally failing to think for themselves about how things could actually go wrong) would probably luck into creating a working cubesat (that isn’t just a carbon copy of some earlier project that worked) like 20% − 40% of the time.
BUT, Eliezer didn’t say “make a cubesat”, lol. That’s like the easist possible space task!! Anyone can make Sputnik 1; the hard part is obviously making the rocket… and then you have to “succeed at landing a space probe”, presumably on Mars or the moon. Yeah this is starting to look completely impossible.
Getting to test the rocket with unlimited retries in the atmosphere actually plausibly gets you most of the way to a working orbital rocket—as far as I’m aware Starship has only done suborbital flights so far, and it’s shaping up as a pretty serious, mostly-finished rocket (albeit these flights have gone into space, but they could’ve done similarish flights that technically stayed in the atmosphere if they had to). In real life of course nobody gets unlimited retries, see here for the assorted failure modes that doomed all the different attempted flights of the Soviet N-1 moon rocket, which flew four times and blew up four times—featuring phrases like:
“One unforeseen flaw was that [the rocket’s command computer] operating frequency, 1000 Hz, happened to perfectly coincide with vibration generated by the propulsion system, and the commanded shutdown of Engine #12 at liftoff was believed to have been caused by pyrotechnic devices opening a valve, which produced a high-frequency oscillation...”
“The engine control system would also be reworked, increasing the number of sensors from 700 to 13,000.”
“One of the largest accidental artificial non-nuclear explosions in history.”
But even if you test everything you can in the atmosphere, your rocket probably still just immediately fails on some aspect of its uppermost stage that’s supposed to push your space probe to the moon/mars. Upper stages are way harder than cubesats: you have to deal with propulsion systems (with valves that can freeze in weird ways in space, and you can’t really “test a propulsion system in vacuum” like you can put a satellite in a vacuum chamber), you have to actually orient and point the proper direction (cubesats can just tumble), you have moving parts like fairings and decouplers that again might possibly behave weirdly in space and are not trivial to test in fully realistic conditions on the ground, most of the burns have to actually fire at the exact right moment and last for the exact right amount of time otherwise you won’t arrive at the moon/mars (versus if you miss a command on a cubesat because it was resetting or whatever, no biggie, just send the command an hour from now when it’s looped around the earth another time). ChatGPT estimates that in the history of rocketry from the 1980s to now, maybe around 60% of genuinely new upper stages have “basically worked perfectly on the first try”—although obviously those were all built by normal non-naive engineers (indeed, the recent wave of move-fast-and-break-things small-launch startups have a significantly lower hit rate than long-established space programs and traditional defense contractors); maybe fully naive engineers have like a 1⁄5 or 1⁄10 chance of achieving similar outcomes, so maybe 6% − 12%.
Building a moon/mars lander instead of a cubesat is a even more difficult than a rocket second stage, I’d say. Once again you are creating a custom propulsion system, lots of commands have to go off exactly on time (ie during landing), you’ve gotta control your probe’s orientation, etc, but now you have this additional problem of dynamically measuring your distance from unmapped rough ground. Any mistakes in terms of thrust direction / timing now have to be corrected instantly or you hit the ground and die, unlike with an in-space burn where small mistakes can probably be fixed hours later with small correction burns. Also, there’s a good chance your naive-engineer’s plan for dealing with space radiation is basically just “YOLO”, so there’s whatever-percent odds that your ship just dies enroute and whatever half-assed reset procedure exists isn’t enough to get it back. And if you’re landing on mars it’s even worse because you additionally have to worry about heat shields and parachutes and maybe dust messing up your distance measurements, who knows. (If you were a non-naive engineer and knew you had plenty of resources but only got one try, you’d be like “supersonic parachutes are too easy to mess up, we’ll just do heat shield + rockets and it’s fine that the probe will therefore be heavier”, but our naive engineers would miss this.) Similarly a non-naive engineer would probably realize “with infinite resources but only one try we should to to extreme lengths to minimize the number of finnicky moving parts like deployable antennas, solar panels, landing legs, etc, which always fail”, but our naive engineers are just going to have to cross their fingers that their solar panels don’t get stuck in some unexpected way (often it’s hard to perfectly test these sorts of mechanisms because the parts are too fragile to work the same way in earth gravity that they would in zero-g). This is probably 3x harder than making an upper/transfer stage, and for our naive engineers let’s say their odds of success on this task are essentially independent from their odds of success on the upper stage task (since in both cases they’re basically just hoping to luck into avoiding various specific potential failures; the whole concept is they lack the kind of mindset that helps them systematically avoid whole swathes of unknown-unknowns failures), so like 2% − 4%.
So overall I would say maybe 0.3% that a smart and well-resourced but suicidally naive team of engineers lands a space probe on the first try.
The contrast between my gloomy estimate of the success probability for the concrete space-probe thought experiment, versus my relatively optimistic vibe in my on-topic AI alignment comment (tl;dr “come on, what is MIRI’s take on these promising-seeming factors that might help AI go well??”) is left deliberarely unresolved as an exercise for interpretation on behalf of the reader.