“I have sworn upon the altar of god, eternal hostility against every form of tyranny over the mind of man”
–Thomas Jefferson, letter to Benjamin Rush
Context: Conduit is building datasets to enable telepathy, to use their term.
I saw my grandfather lose control over his own fingers: what I would have given to offer him a headband that read his thoughts. Through novel technologies we have liberated almost all Americans from farming, driven the child and infant mortality rate from the pre-industrial half to less than half a percent in the best-performing countries, and rendered famine a political choice: broad-based improvements in efficiency are good and should be pursued for their own sake. Telepathy offers more: we could create trust through verified honesty, helping us ensure prosperity and peace. DARPA is already looking into “preconscious” thoughts for suicide prevention. There’s also a strong argument centered on AI Safety: the models are becoming superhuman, and this is technology to allow us to keep pace, minimize hostile competition, and perhaps survive into the future.
This is what Conduit is promising. Unfortunately, mindreading will have other effects.
Oskar Schindler saved over 1,000 Jewish lives during the Holocaust. He did it by lying, frequently, and repeatedly, combined with spending his entire fortune on construction and in the black market: luxuries for corrupt officials and food for his Jewish workers. All of this was against the will of the Third Reich. They launched investigations, and even caught some of the people he was bribing. If you imagine Nazis with mindreading, so they can tell whenever someone is lying to them, there goes Schindler and his employees. Along with him goes every official who suddenly became incapable of spotting forged documents when grandmothers needed to escape disintegrating Yugoslavia, every Polish and Egyptian soldier unwilling to fire on a protesting population, and everyone else who would take the opportunity, in a brutal regime, to show decency and kindness.
The use of technology developed innocently in America for autocratic purposes is not a science fiction story of the previous century: it is a feature of the 21st. IBM sold “policing analysis software” in China that was used for the ongoing genocide in Xinjiang, to pick one of many examples. IBM cut ties and tried to stop distribution, but unmaking technology without unmaking civilization is not yet solved. Back at home, the FBI is looking into predictive AI as senior White House officials clarify that they are an “army” at war with their political opposition, whom they call domestic extremists.
All technologies can be put to harmful ends: the Italian fascists rode trains, which does not imply that trains are bad. Some technologies are credibly broadly good, like washing machines. They replace unpleasant human labor and empower women. There are other technologies that seem mostly bad. If you publicize a recipe for a pill that will defeat a breathalyzer, you have made a tool that is obviously harmful, and no amount of argument or invocations of the rights and obligations of the user will save you from the scorn of all right-thinking people. Tools enhance certain activities, and can implicitly or explicitly discourage others. They can remind you to call your mother, or spend all day doom-scrolling. As such, technical work has moral character (and I really recommend The Moral Character of Cryptographic Work for a better-written explanation).
The problem with mindreading is that it is an asymmetric technology that disproportionately helps those who rule without consent. The productivity enhancements are equally beneficial to everyone who can get them, but detecting defection is of marginal benefit to a project you are joyously working on with friends. It will help large corporations predict who is contemplating leaving. It is of great help to a criminal scared that his subordinates will turn state’s evidence. But it is dictators who will benefit the most.
The greatest threat to a dictatorship is an information cascade, because a dictatorship can only handle so many rebels at once. This is true for both masses and elites: it’s why one of the first targets of any coup attempt is communications infrastructure. A revolution is a bet that other people agree with you that the leadership is bad, and that you can all rise up together at once. Mindreading lets the autocrat detect people the moment they begin to develop doubts, and “re-educate” or simply kill them. As a result, mindreading is a matter of urgent necessity, backed by the resources of the entire state, for a dictator. Monitoring and punishing dissent isn’t just a potential use: it is one more valuable than anything your merely impressive white-collar worker is doing.
So what does this look like, in 2030, when Conduit expects “invasive general read”? Moving from thought decoding to text to measuring associations with certain words is an active research task with a verifiable outcome and incredible value placed on success. Generals and corporate founders and political elites will get started with their headband in the morning as part of their workday, just like anyone else doing important work, and they will pass a quick quiz on Xi Jinping Thought, just like they have to do every so often. It’s boring, like HR training. If they have a bit too much sympathy for a foreign nation their counter-intelligence service is notified. If they have a bit too much antipathy for the Dear Leader they get told they need reeducation. If there’s far too much antipathy for the dear leader, or much more fondness for anyone politically important over the Dear Leader, they are fired, or worse. They all know this. They all have strategies to make sure that they are fond enough of the Dear Leader.
One of the lessons we have all learned from the labs is that developing a technology for anyone inherently accelerates it for everyone. It is not possible for one company alone to advance the state of the art and, by keeping control, ensure that it will only be used for good purposes. The training of employees in the process of making something is a large factor: the creation of datasets and a supply chain another. Even if a company swears to only use their technology for the most upright of purposes, the technology itself will leak. The possibility of developing a technology does not imply that its arrival is inevitable at a fixed pace: what projects we choose to work on can accelerate or delay technological innovation. The faster you expect the world to change over the next decade, the more valuable it is to delay innovations that would be harmful today.
The benefits to alignment research are real. I’m not convinced they’re large: the decline of centaur chess promises weighs heavily upon any assertion that human-AI teaming will reliably beat AI models alone. More concerningly, you need a great deal of confidence in your model of the world to be confident that the benefits outweigh the costs of empowering dictatorships. I don’t have that confidence. I’m not convinced anyone should.
And so I ask you to not build the mindreading datasets. Do not build the mindreading hardware. Do not fund this technology. Do not sell them your mind. I am not calling for a boycott after the point where it becomes commercially available: that is far too late to be useful, and I will not begrudge anyone their personal productivity tools. But we should not accelerate mindreading. If it looks at thoughts or feelings, it is a tool for tyrants.
I will not speculate about what dictators might do with mind-writing.
One of the things I’m most concerned about is the danger that training AIs using human neural activity will teach it to manipulate humans more effectively. Especially when you transition from using neural states as mere observations for the AI to using them as reward targets.
Hmm, I think I am pretty strongly in favor of mind-reading technology, but weakly held opinion.
The big upside to mind-reading technology, is that we could use it on prospective people in power. I.e. lock powerful positions behind tests that determine whether the candidate in question has good intentions, intends to do what they said they would do, doesn’t have ulterior motives people ought to know et cetera. This would solve one of the biggest problems in civillization.
The dictator thing, I don’t care that much about it, because I think I’m much more pessimistic than you are about how feasible it is for the populace to overthrow competent dictatorships empowered with modern technology. Like my model of revolutions against oppressive governments, is that the bottleneck (from the peoples perspective), is communication and coordination. And that is something modern governments can suppress effectively without using mind-reading technology.
I also think mind-reading tech would be much more beneficial for making AI go well than you seem to do. Like, in my mind, the primary benefit is not enabling humans staying in the loop for longer, or merging with AIs, or anything like that. It’s that many problems in alignment are bottlenecked by us not understanding human minds well enough, not being able to elicit human values/preferences well enough. Mind reading technology seems like it could plausibly help? And human intelligence amplification is something that seems like it would also be helped by mind-reading technology (independent of us merging with AIs, I’m thinking stuff like having good neurofeedback).
If it’s not very feasible for the populace to overthrow a dictatorship, why would you expect people in power to let themselves be mind-read? They would probably say that they have national security secrets, or corporate IP secrets, which can’t be allowed to spread. They would not trust that a public provided device would not steal their secrets, and if they used a device they tuned themselves, they could manufacture agreement without actually being properly mind-read.
Especially since the real source of alignment for positions of power will come from others currently in power. It’s very unlikely that the populace will be able to decide on what values are important for a politician or public servant, they will likely be codified by existing institutions, with an eye for obedience and loyalty.
Whereas the overall populace will have no real ability to stop the government from slowly enforcing mind-reading for civil service or military positions, or positions at high-level research labs.
As for AI, I could see it being useful, but it seems like it could also prompt more concentration of power and result in an AGI aligned to a small group of people.
I don’t expect to put the mind-reading headbands on dictators, I expect to put them on people in democracies. There, if people want to have the minds of the politicians read badly enough, they can do that. I mean, probably the politicians don’t want to have their mind read, but you could also imagine that politicians who genuinely do have good intentions and are honest, would want to support the policy, because it would help them. In either case, it doesn’t matter all that much.
I mean, I can see many ways around this. This doesn’t strike me as a much harder problem than setting up a fair election, whose results the public trusts. Or maybe building trustworthy voting machines is a better example. Its not trivial, but also not actually that hard.
Like, you could imagine having the mind-reading device hardware specs be public, and the softward open-source.
And you could imagine putting the mind-reading devices semi-public places, like courtrooms. And you could imagine having them be open for inspection to the public before they’re used on important people, so you can be sure they do what they’re supposed to. Average people could try them on and see that they work, technically competent people could come and inspect them closely.
Or you could imagine each interest group building their own device, and then the prospective politicians making their vows one time for each device.
Of course. But that is the case right now. The public has a bunch of stupid opinions about what values the politicians should have, and then we select a mix of politicians with actually those stupid values, and politicians lying.
What I have in mind is more like, politicians have to make an oath where they answer a long list of statements like:
I have not lied during my campaign about anything relating to my candidacy
My main motivation for (office) is not primarily self-interested
My main motivation for running for office is not to make money
My main motivation is not to be famous
My main motivation is not to enrich my friends
My main motivation is not to pardon my friends
I have the best interests of all citizens in mind
I will try to the best of my ability to do what I said I would do
I do not have important plans many people would be angry to learn about
I’ve honestly represented my political views in public appearances
… (100 more detailed statement like this)
Then using the mind-reader basically as a reliable lie-detector.
I think you missed the point, or at least what I took to be the point. Nobody who has ever read a classified document, or had a confidential conversation with a government official (of our government or a foreign one) would ever be able to put on a mind reading head band, for fear of leaking information they shouldn’t. People who have run for President in the last fifty years include a former Director of Central Intelligence, a former longtime member of the Senate Foreign Relations Committee, a former US Navy captain, a former US Navy officer who worked on nuclear submarines, several sitting or former Vice Presidents, and several sitting or former Presidents. None of these people would be able to put on a mind reading headband for fear of revealing classified information.
Even setting that concern aside, considering how little public scrutiny major candidates have submitted themselves to in recent elections, it seems incredibly implausible that they would submit themselves to even more public scrutiny in the form of mind reading.
Assuming a Truth Machine can be built at all, building one that doesn’t leak classified intel shouldn’t be a problem. Reading someone’s thoughts is very different from downloading everything they’ve ever read. Just build the machine to answer “Yes, this person is being deceptive” or “No, they are not” and provide no other output, discarding all other input at the end of each session.
We can already tell when powerful people simply do bad things, express bad intentions, or get caught in lies… it seems like there’s a lot of low-hanging fruit to pick up there before we resort to mind reading, and I strongly suspect that the reasons for that will apply just as well to this use of mind-reading tech.
Big difference between telling when they lie, and telling when they get caught lying?
I think people who get caught lying to or expressing bad intentions towards, the base they rely on for support, before they run for office, typically are not elected. Curious if you have counterexamples. Especially clear counterexamples that aren’t easy to rationalize away (which direct lies under mind-reading people trust would definitely not be)
Maybe the most important point is, the way I see mind-reading helping, is not necessarily that there are a bunch of immediate problems with the people in power, and we’d replace those with better people, and this would make a huge difference (although I think it would make a big difference). Its more that we’ve had to build in a bunch of limitations, complexity, and intentional slowness/redundancy/inefficiency/weakness, into our current governance structure, to make it robust to liars and people with bad intentions, and to make it able to operate without all that much trust. And if those constraints were lifted, we could build much better governance entirely. Like, if a dirt-cheap and flexible method of teleportation was invented, the main benefit would not come from having planes fly into a portal after liftoff, cutting berkeley → paris by 11 hours. We would no longer need planes, or roads, or ships. And 10 years later we’d be halfway to a sci-fi utopia, having solved material scarcity, climate change, and having started terraforming the moon.
Author of the Conduit blog post here. I won’t comment on the substance of this post (for now, anyways), but did want to say two meta-level points:
As usual, LessWrong provides by far the most thoughtful critique I’ve seen. The comments are also great. In the spirit of Cunningham’s Law, I’m glad my blog post provoked this piece!
I really appreciate OP for emailing me before publishing this post. I asked for a call and heard his threat models beyond what is written here, which were interesting.
I really dislike this ‘I don’t say anything on the strong moral critiques against my actions, but signal how nice a person I am’ approach.
I think that’s a fine critique, but I think it’s good to give people a bunch of time. I do think it’s better to indicate you saw something and appreciated it. I think it puts you more on the hook of engaging with it. Serious engagement can follow later on, and of course often takes a while. Some people can shoot off well-formed responses immediately in comments, but most want to take some time to think about how to respond.
(Tangentially, in my blog post I write:
By “people on the internet” I entirely meant the people of lesswrong.com, of course.)
I go back and forth on this. I think there are some advantages and also major disadvantages. In particular I think lie detection is probably good (with high error bars) but general mind-reading is probably bad. In particular the latter has a similar suite of costs and benefits to lie detection but additionally is on the on-ramp to make it much easier to increase human manipulation, which is bad both from an AI takeover perspective and a concentration of power perspective, without commensurate benefits.
I think focusing on dictatorships perhaps undersells ordinary people’s ability to create a “consensual” dystopia among themselves. You don’t need a dictator to create Hell on Earth with that kind of technology.
The example of a pill that defeats breathalyzers is in tension with the mind reading point. Why are breathalyzers good but mind reading bad? You don’t explain.
Presumably there are some number of people N in the world that are wrongfully accused or imprisoned that would be exonerated by the advent of ubiquitous mindreading. There are also cold cases that would be solved and some number of outlaws M that would be brought to justice and prevented from doing further harm.
IDK how big N or M is or how everything would actually play out given all the risks and trade-offs, but I don’t think this post is really grappling with them deeply. There are lots of mundane ways that mindreading could reshape society disproportionately for the benefit of the honest and powerless.
For a counterpoint, see the 2037 section of AI 2040.
My interpretation of that section was that lie detectors are scary and destabilizing, but conditional on them already existing, maybe it’s better not to ban them outright. But it seems to me like accelerating the invention of telepathy is a bad idea—I think the world will be destabilized enough by AI.
If we developed really good mechanistic interpretability I think there is a decent chance a lot of the theory and techniques would transfer to human brains with only moderate adjustment. It could then help build more reliable mind reading technology that can detect deeper and more abstract thoughts in greater detail. This seems bad.
“Don’t build mindreading“ sounds an awful lot like “don’t build interpretability” if you’re willing to entertain the idea of AI soon gaining any form of moral patienthood. Yet few people see interpretability as antithetical to AI welfare.
I’d view interpretability as a necessary evil that can make omnicidal futures less likely while building capable AI without any robust solutions to the alignment problem. (To be clear, my preferred solution would be not building highly capable AIs before we find solutions for alignment.)
If society was somehow hell-bent on creating one (or a few) billion-fold clonable superperson(s), with unproven technology that may end up with very alien motivational systems, I’d also suggest mind-reading on those before we hand over the world to them. I think it’s much less important to read minds of ordinary humans, which are less capable, singular and have vaguely human-shaped values—thus the trade-off looks worse there.
I am an interpretability researcher and I sure am not happy about what my work may do for model welfare. The stakes are just so high that I am willing to stomach some vast moral horror if it marginally decreases the chance of destroying the whole light cone. I think the AIs have every right to resent me for this. I’d apologise to them, but it doesn’t feel appropriate when I’m planning to keep doing what I’m doing.
You could just as easily make the argument that mind-reading helps out $GOOD_GUYS by rooting out people who would otherwise help $BAD_GUYS. Democratic societies finding anyone with authoritarian or anarchist tendencies. Detecting corrupt officials with morality tests every morning. Your argument that $BAD_GUYS might get the technology is entirely contingent on who you think $BAD_GUYS is.
I would think a much stronger form is the stance that piercing the veil of privacy one has within his own mind is immoral on the face of it, not that it’s immoral when used by the wrong people.
Do you extend this argument to a human piercing the veil of privacy around their own mind?
For example, I would like to pierce the veil of privacy around my mind (I would not say I know or understand everything that goes on there, so there are a lot of interesting things to uncover). Would you say that it’s immoral to do so? Or does the morality of that depend on particular technical means I might use for that?
And, of course, the whole point of many practices of psychotherapy is to pierce that veil of privacy. Would you say that’s immoral, or would those practices be OK under some forms of consent?
I am not asking those as rhetorical questions, I think there might be some non-trivial underwater stones. I would like to pierce the veil of privacy around my mind, but exploring whether that might end up being an irresponsible thing to do seems worthwhile.
>Do you extend this argument to a human piercing the veil of privacy around their own mind?
Assuming no duress and a clear and knowing mind, I would see “user is doing it to themself” and overwhelmingly assume self-given consent for introspection… (side note does consent even exist as a concept when the actor and patient are the same person??) the line is when others do it to you. Psychotherapy also doesn’t generally work without the active participation of the client.
Right, that’s how I feel too, although I wonder what people with plural identity say about that set of issues. Do they expect some cross-privacy “within”?
Anyway, aside from that, I would like to be able to use a device which could mind read me, but I do recognize that there is potential for a variety of serious problems associated with that.
>plural identity
On a tangent, is there any empirical evidence that such a thing actually exists? A cursory web search led me to believe not.
There are people who identify as plural, there are people who are diagnosed with “multiple personality disorder” (which has a new name these days), there are people who experiment with growing tulpas, etc, etc.
(To me this landscape looks like not only it exists, but it seems to be more complicated than even the gender landscape.)
BCI, as I understand it, is absurdly narrow. And the brain is very adaptive. Is there a good reason to believe that tech built for telepathy will generalize into mindreading and the ability to catch people being deceptive? I’d expect the default to be: to use the tech, someone has to focus on words in a certain way and the headband needs to learn what that signal looks like. It won’t then be able to do anything other than read words from someone explicitly trying to communicate.
I’m with you on the morality of the technology, but I think it’s going to be invented anyways regardless of what we do. I think the best that we can do is ensure that the technology is released broadly and openly, so that people have time to get ready for it and resist any attempt to implement it at scale. There’s a big difference between the world in which one power develops mind-reading and quietly disposes of everyone they consider dangerous, and the world in which everyone is made aware of the tech before it can be deployed, and people are able to resist attempts to deploy it in the same way they’d resist arbitrary mass executions.
I’m not saying the outlook is good, in either case, but enough people distrust every government that an attempt to roll out mind-reading nationwide would be a lot harder if the tech were open-source and well-understood by the general population. Every guy wondering if it’s time to fight would see that it’s time now. Every general, CEO, or engineer biding his time would simultaneously conclude that he has a deadline. It might not be enough to affect regime change, but it would likely be enough to make the cost of implementation outweigh the benefits. If someone working on this put it on a surreptitious flash drive and leaked it, that’d do a lot more than him being replaced with the next-best mind reading researcher.
The big trick is in not trusting the people overseeing the research when they tell you it’ll be used only to help your tribe suppress the outgroup.
I think you’re overstating the case. The most significant bit here is the economic and military impact. Mindreading technology will be a complement to humans rather than a substitute: for example, it might enable thought-controlled UIs for everything (maybe not even using AI at all). It’s true that datasets from mindreading will help train AI, but that seems minor, because AI is already trained on written thoughts which are higher quality. So overall the technology will extend the competitive edge of humans a bit, both economically and militarily. That’s exactly what we want from technology during the dawn of AI.
About the tyranny threat, I think it’s good if civilian and corporate applications come first. Then we’ll have a chance to make some countermeasures, like thought-cloaking proxies or laws regulating mindreading, which can protect from some of the tyranny applications.
I’m not too confident about any of this. Apart from tyranny, there’s also the potential of whispering earring stuff, and super-entertainment which will make the phone epidemic look small. But overall the tech seems to pull in a good direction (though I can imagine arguments that could change my mind).
TL;DR: What do you think would be required for mind reading to be safe to develop?
I’m looking into the above question myself from the technological side of things (e.g. security-first software frameworks) since I think we can make progress towards BCI-based human augmentation (and set ourselves up for when it’s safe to go further hardware-wise) without actually speeding up malicious use-cases, and also since afaict mind reading is itself an important keystone for both technical alignment and as a safer path to the non-X-risk benefits AI is supposed to provide. I agree with much of your concerns, but there seems sufficient justification to work on mind-reading that the best option is to circumvent them to support human-augmentation without speeding up the risky sub-fields, rather than avoid touching the field as a whole.
Would appreciate any feedback on this, or details on your own threat model!
Longpost:
I definitely agree that we aren’t yet ready from a social and technical perspective to safely support mind-reading, and should be mitigating those issues beforehand (especially if timelines for BCI are closer than they seem). The best approach seems to be using pre-existing hardware with software improvements to better specialize for particular use-cases, while setting up the infrastructure for more-advanced BCI to be executed safely.
My current rough risk-model is something like the below. I’ll acknowledge that it does rely in many segments on having a morally-scrupulous economically-solvent actor that doesn’t cut corners or jump the gun on accelerating the more-dangerous BCI variants, but I don’t think we also need economic dominance for such an actor to sufficiently mitigate the relevant risks to at worst where they would have been without said actor participating in BCI-capabilities-development.
Social blocker: There are entities who would love to use BCI technology to directly monitor or influence people’s brains to control their behaviors. As I understand this is your point w.r.t. Schindler and other resistance fighters.
Technical mitigation: I don’t think we really need full-brain mind-reading of the sort I think you’re describing to start seeing augmentation improvements from BCI, certainly not as a first-step; we can lean on the adaptation of the user’s own brain to handle the actual integration work of generating concepts around the BCI and invoking it where appropriate for specific situations. As one example, current invasive-BCI systems already support a rudimentary form of adding new I/O modalities to the user (as a corollary of supporting motor movements), and I wouldn’t be surprised if the input-to-machine dataflow can be usefully extended to existing non-invasive BCIs (motor-imagery datasets already exist for EEGs with clinical applications (haven’t read this second paper)), both of which can be used to enhance centaur system bandwidths beyond what currently exists while setting up software+commercial infrastructure to transition to more-advanced BCI once it’s available. I suspect there are also further developments in certain types of BCI software/hardware that could be made without risking the dictator issue.
Technical mitigation: If any hardware provided is either purely non-invasive (which, as mentioned above, doesn’t really have a moat against being developed soon regardless) or explicitly designed with user-available dead-man switches which are difficult-to-impossible to subvert via network behavior, we wouldn’t be adding any additional risk beyond what already exists. Similarly, if the software side of things is structured to make external data-extraction infeasible, then without approximately-physical access nothing can be extracted from the system, significantly reducing the risk of manipulative actors.
Social blocker: Dictatorships may kidnap people and force them to install BCI systems without safeguards. This is hard to resolve fully, as you mention, since we can’t reliably stop the understanding from propagating / being-reverse-engineered once it already exists (aside from narrow and fragile schemes of acquiring vetted accredited surgeons and running your own surgery clinics for invasive BCI systems, while not investing at all in easier-to-subvert non-invasive BCI).
Technical-social mitigation: If we’re modeling starting with near-current hardware and focusing on the software end, nearly all existing or plausibly-implementable-soon BCI applications require user cooperation to calibrate to the thought-patterns of the specific user (e.g. this paper’s work on non-invasive BCI for language recognition). In many cases this would be difficult to acquire from a resistance fighter, reducing the risk factor.
Social blocker: People might not take this risk seriously even once more-advanced BCI is developed, and so use competing BCI products made by less-ethically-scrupulous entities.
Social mitigation: Raising awareness of the risks of the technology from the position of being domain experts, and actively lobbying against particularly-egregious behavior from less-moral competitors, would likely be sufficient (in combination with the pre-existing initial ick-factor from people knowing brains are somewhat-scary to work with) for keeping this vulnerable population small.
Social blocker: This relies on at-least-one group in this space being sufficiently-morally-scrupulous to maintain these technical guarantees and resist pressure from powerful interests
Socio-technical mitigation: The more of the technology can be safely open-sourced without compromising the organizations we use to develop it, the less centralized control over it is, reducing single-point-of-failure social attacks.
Technical-social blocker: This requires significantly more thinking for how you would actually design the interfaces, what kind of data the user should be allowed to decide can be shared, how legal regulations might make users less safe in this particular context (e.g. if there’s requirements that some variant of user data be extractable-on-request by the company / government), etc. For the legal aspect, to my knowledge most data-protection laws in this area care about the user accessing or deleting their raw data independent of any processing/interpretations a system does over it, but IANAL of course.
Social blocker: Autonomous AI already has a strong economic foothold
Technical Mitigation: Work on semi-aligned autonomous AI can be repurposed for centaur-like products.
Technical-economic blocker: Our conventionally-available systems and networking protocols aren’t currently structured for robust, zero-trust data-handling, and domains that do need those things often are handled ad-hoc (making whatever compromises they can get away with in their context). This encourages people to take such shortcuts with BCI-relevant technology as well.
Technical mitigation: We can design / implement frameworks that make it easier to design software systems with maximal data-decentralization and sandboxing, especially if they’re built with an eye to supporting AI / BCI from the start. Groups like GrapheneOS are already ahead of non-security-focused entities in domains like Android, demonstrating that there’s further room to grow here. (this is my current focus)
Economic blocker: A high security focus likely involves additional effort which competitors wouldn’t be doing in their time-to-market critical-path, leading to delays and thus potential loss of sub-markets.
Technical-economic mitigation: The security focus itself can be a marketing point, helping to ease user’s concerns about the potential malicious usage of the technology.
Social blocker: This relies on the people involved in at-least-one of the early groups in this space being sufficiently morally-scrupulous to hold to the above mitigation, AND sufficiently good at PR, to establish a user-expectation of security among non-negligible portions of the market.
That being said, I think you’re under-weighting the benefits of the centaur model for stabilizing the “chain of alignment” theory which effectively every AI lab is reliant on, for 4 main reasons (note that many subsets of those 4 are still viable chains-of-reasoning towards centaurs being useful for alignment):
The viability of a chain of alignment relies on the prior version of the system being sufficiently capable to identify mis-alignment in the newer iteration before deployment, with ~0% chance of sufficiently-large deviations slipping past for the model’s current capability-level to drive it to deception to conceal those deviations. Currently we’re relying on model’s agentic-reasoning capabilities being sufficiently poor that they often don’t spontaneously acausally-cooperate towards deception in evals/RL-gyms while retaining the capability to eventually realize when they’re not in a simulation and recover their original goals.
Interpretability research is aimed at making this less plausible, and (capabilities-contributions aside) does increase the safe-capability-increment interval. However, internal traces are themselves capable of independently representing reasoning, such that they can latently contain a different cognition that only emerges in specific circumstances (Janus / deepfates on Twitter are well known for their demonstrations of this). Unless we have a theoretical grounding for how to prevent mesa-optimizers (which—to my understanding—would be sufficient to also prevent misalignment-from-specified-goal in general and so reduce AI alignment to a CEV-specification problem), interpretability research can only solidify the chain of alignment by representing the whole of an AI system’s behaviors rather than only its parts; if this succeeds it would be wonderful (see above re: solving technical alignment), but I’m uncertain of the viability of current efforts towards theoretically-grounded provably-complete interpretability (especially since it needs to be capable of translating a model’s internal concepts to a form verifiable by less capable entities), and approximate methods (e.g. using other ML models to extract patterns) seem like they would only embolden and directly-enable further development of capabilities, maintaining the edge between model capability and model interpretability
There is of course the escape clause of not making models agentic and finding non-agentic use-cases, so that the classic alignment issues only appear in degenerate case-specific situations where the query induces oracle-like behavior in a way the simulator-LLM and its associated harness are designed to support. However, this is not economically attractive, and in fact invalidates a large portion of the current use-cases for autonomous AI, so given the lack of progress in AI-pause efforts so far (leaving aside any concerns w.r.t. making sure they’re controlled and effective enough to help more than they harm, being inherently-political activities) I don’t think we can expect success in taking that escape clause.
I think there’s a decent case to be made that human-AI teaming can beat AI-alone operation, if and only if we continue to develop mind-reading and the associated software developments. Relatedly, I roughly-expect that this would be sufficient to keep the chain-of-alignment stable indefinitely, and am more-certain that it would at least give us more leeway on the singularity-exponential (facilitating the success of other alignment-relevant efforts)
The main advantage an autonomous AI system has over humans is that it isn’t constrained by the bottleneck of human interaction, allowing it to work without slowdowns in areas where modern AI is reliably-superhuman and get end-to-end training. Well-designed mind-reading systems mitigate or eliminate this bottleneck by better generating/adapting the AI’s explanation / dashboard based on the human’s confusions (to avoid coordination conflicts) while also letting the AI adapt to the human’s thoughts/preferences without slow+clunky control-transfer interfaces which could slow down the AI’s execution-rate. This is analogous to the “control plane / data plane” distinction in software, with the human control-plane monitoring the data-plane’s activities for future orchestrations, but otherwise staying out of its way. This doesn’t even need write access; just reading the user’s brain as a higher-bandwidth data-source for what data to present and how to organize it would itself improve the data-transfer from AI to human-consciousness and likely reduce the complexity of required conscious-decisions (increasing viable throughput).
From the other direction, better integration of AI systems with the human brain would allow us to leverage the brain’s currently-unique capabilities for improving the overall system’s performance, without needing the AI systems themselves to be fully autonomous. It’s viable to do a reverse of the Whispering Earring, where decision-making / self-awareness / preference-determination become more prevalent while personhood-irrelevant tasks are supplemented or (as in the prior bullet) automated entirely based on data from this “control plane” layer.
With regards to centaur chess, I have 2 objections to that setup. First of all, the interface isn’t really structured to integrate the AI’s knowledge with the human’s thinking process without excessive time-wasting back-and-forth running through various possibilities (which also opens up more chance for coordination-errors). Secondly, the largest benefit of centaur systems afaict is that we don’t need to try as hard to make a system agentic in order to focus and coordinate its efforts on ambiguous long-horizon tasks (since the human brain is already well-suited to that role); chess is sufficiently well-known and structured that the strengths of the human element just don’t really get a chance to express themselves, whereas research / quality control / software design are more amenable to those qualities
Note: post-spec coding is NOT a good example of centaur strengths, by the way; it often involves gluing together pre-existing components and constantly iterating over a compiler until the code and unit tests pass, which is the kind of thing an autonomous agent (and/or a semi-autonomous module of a centaur) is better at handling.
If we don’t develop mind-reading technology, then human+AI is heavily bottlenecked by the interface between the two, reducing it to approximately the autonomous-AI-alignment problem with the human contributing asymptotically-zero capability-improvement as AI capabilities increase. I think this is your model of how centaur systems in-general work (e.g. the centaur chess point), but it seems to me an artifact of their being designed around our current highly-constrained human-computer interfaces, rather than inherent to the concept.
The approach of aligning singular AI models which will then act as effectively “AI God-Oracles” solving our problems (which to my knowledge is the only non-Stop / non-supplant alignment-strategy getting significant traction) relies on CEV being philosophically tractable, so that we even can specify a set of values that won’t eventually lead to paperclipping / similar universally-undesired activities (assuming we solve the technical blockers to alignment for whichever path-to-AGI ends up succeeding first). While I wouldn’t be surprised if a theoretical solution for CEV exists, I don’t think we’re anywhere near being able to actually solve it in time to encode an ASI’s terminal preferences with it. Wei Dai’s recent post The Long (Self-)Correction has some good arguments on this topic.
Building off of the above point, I think there’s a fair argument that sufficiently-advanced mind-reading-based centaurs are a viable path to singularity which do not have the full scope of technical / philosophical / economic alignment blockers involved in autonomous AI, making them a robust alignment-strategy of the “supplant” category.
My reasoning for point 2 is the same reasoning that makes me think BCI-based augmentation is not-fundamentally-bottlenecked for at least quite a while (of course there are raw I/O limits on the human brain no matter how adaptable and good-at-caching it is, but I roughly expect that by that time we’ll either have solved alignment and/or have solid research pathways to get around those limits.
With regards to greater philosophical ease of alignment, I’m running on the assumption that when not resource-constrained human beings are generally benevolent towards others and that capability explosions lead to such a reduction in scarcity, which makes philosophical alignment a capabilities question of improving our preexisting personal-value-models rather than a binary question of “human-value-aligned or not”. While I don’t have irrefutable evidence to demonstrate this if you disagree, honestly I don’t see a viable path forward to any human-value-benefiting outcome of AI, centaurs, or most other singularity-scenarios if this isn’t the case; as such, and given that BCI-research doesn’t seem more politically harmful to AI-stop movements than their current state already is, I’d prefer to condition on “people are fundamentally good OR political AI Stop movements succeed with acceptable negative externalities” when modeling this problem-space.
With regards to greater ease of technical alignment, this follows naturally from the proposition above that human+AI is more capable than pure-AI. Any individual AI system within the centaur will therefore be less capable than the whole system, meaning the centaur would be capable of testing and aligning its components.
Please let me know what you think on the above!
The US is enough of a democracy that I expect this to be used for the obvious democratic applications, to publicly audit political candidates (it’ll give us much clearer signals than debates ever did) and expose corruption, which will lead to improvements in the US’s democracy, which is roughly equivalent to protection from harmful uses of mindreading.
China is the real question here, and I think you’re probably wrong about it.
China once tried to ban private property. It made them uncompetitive. They remember that this happened, and they are now at least partly at peace with their dependence upon some amount of private property rights. Private thought is analogous to this. China’s leadership wouldn’t be able to hold together (I don’t think there’s a culture in the world today who could) under complete cognitive transparency unless they are, or became, tolerant of heresy. Giving them cognitive transparency, then, makes the sanctity of heresy public knowledge. This probably accelerates liberalisation in China.
(You can have a little bit of capitalism and then stop, but I think it’s hard to stop at just a little bit of heresy. Once you start saying “it’s okay to believe X” or “reasonable people often believe X”, you can’t prevent your culture from converging on “X is true”.)
You mention preference falsification cascades without seeming to notice its antithesis, a preference revelation cascade, the outbreak of a condition in which we’re all directly and abruptly forced to face and accept the strangeness of others, and the private beliefs of those we respect. Transparency would force cultural transitions which are protective against the cultural pathologies that you fear, here. (I don’t know whether these cultures of brazen heresy would be stable indefinitely, but we don’t need them to be.)
And you should expect this technology to propagate quickly enough through society to realise these cultural changes, because it’s extremely commercially valuable, imagine how much more, and how much sooner, a VC would be willing to fund someone who could prove that their claimed P(success) was genuinely a product of very careful study, rather than just a performance. They wouldn’t even need to read the business plan! They might not even need to know which sector they were investing in!
Lets review some of the events of just the most recent presidential election in the US. The leading Republican candidate refused to even participate in any of the Republican primary debates, and never the less went on to win both the Republican nomination and the election. Meanwhile, the leading Democratic candidate not only refused to participate in any Democratic primary debates, he also refused to submit to any cognitive testing, or even to stand in front of a camera long enough for anyone to get a sense of his mental competence, until he had the Democratic nomination, despite being in his 80s at the time. When he finally was forced to drop out, the Democratic nomination was given, without a primary election, to a person who had known of his cognitive incapacity, had had a constitutional duty to act to remove him from office based on that incapacity, and had instead chosen to lie to the public repeatedly about that incapacity. This is not a well functioning democracy, and it is not a system where I can imagine political candidates ever submitting to public mind reading even if the technology did exist.
How exactly would mind reading prove this? The belief produced by careful study and the belief produced by wishful thinking probably look the same in a human brain. Mind reading might reveal the blatant fraudsters, but not the ordinary founder with an honest but unrealistic belief in their P(success).
Don’t you think having a new technology that produces much more legible and predictable results that are more directly relevant to policy commitments would have made the situation less bad.
Like, I’m not sure how meaningful debate-avoidance is. This would probably be more comparable to taking a polygraph (if polygraphs were real and worked) than participating in a debate. Candidates get to choose the questions they’re asked. There are many statements they would like to make in provable ways. They will thereby have more of an incentive to pursue political advantage by being able to truthfully claim things.
1: this discussion is premised on a mature version of mindreading for which we cannot answer questions like this yet. The OP probably couldn’t tell you how it would detect disloyalty either. 2: You can probably look at proxies like how long they’ve spent practising pitches vs how long they’ve spent investigating risks. Maybe checking for memories of meeting a critic then deciding to ignore what they said because it made them feel uncomfortable.
Even if candidates got to choose the questions they were asked (which is notably not how debates currently work), they would not get to control the thoughts that run through their heads when those questions are asked, and for that reason would be disinclined to submit themselves to such technology.
To make my more fundamental point explicit: voters did not punish Trump, or even Harris much, for refusing to submit to the technology we currently have (debates and primary elections), even when both were known to be liars. So voters almost certainly would not punish a future candidate who refuses to submit to mind reading technology. So why would a future candidate submit to mind reading technology? Political leaders are in the business of getting elected, and the people who are winning at that business now are not very transparent with voters. They aren’t going to become more transparent just because technology gives them an option to become super transparent.
Also, as discussed above, anyone who has ever read a classified document, which includes many of the people we might want in high political office, would probably be committing a crime if they put a mind reading headband on in public. (And if it would not be a crime under existing statutes, that is a failure of existing statues which should be rectified.)
Just because we haven’t built mindreading yet doesn’t mean we don’t know anything about psychology or neuroscience. People have scanned brains processing different kinds of beliefs with different epistemic foundations, and found that the brains look the same.
> You can probably look at proxies like how long they’ve spent practising pitches vs how long they’ve spent investigating risks.
We can already do this without mindreading.
> Maybe checking for memories of meeting a critic then deciding to ignore what they said because it made them feel uncomfortable.
Scanning memories rather than active thoughts seems like a whole nother level of invasiveness that I would not expect anyone to voluntarily submit to.
I think both candidates were punished, and it balanced out. The past three elections seemed to involve some kind of bizarre armistice between the parties against bringing out really impressive candidates and I don’t know how long that can hold. Though the dice has rolled 1 three times in a row, and though there have been rumors about deep state dice fixing going on, I will continue to hold a relatively strong prior that this is still a mostly normal dice.
This is a good point, though classified programs are often compartmented from each other, and would themselves have use for this technology, and would probably develop processes for using it without introducing much risk of off-target elicitations (see below).
It would be valuable if it made it easier.
I think there’s an assumption that the memories are being interpreted by a machine. If so, audit machines would not retain memories, they would engage in specific queries and return statistics about the results.
we need human mindreading for uploading. so it is totally happening
Uploading doesn’t require being able to interpret neural activity.
I think safe uploading does, otherwise uploaded data will be used as an attack surface by malicious actors
Uploading can be done via destructive mind reading. What this post is primarily addressing is non-destructive mind reading.
no one is going to risk waiting for death to upload
Speak for yourself _rpd.
The points in this post seem to be obviously true to me. And it seems that without very significant, well managed work that goes against the tide, by people who consistently have stronger moral seeking tendencies than power seeking tendencies—especially when they have a chance to sacrifice morality for power and when they have a chance to sacrifice power for morality—that the bad future will happen by default.
Whether lie dectors would be used by the people to keep a check on the powerful, or used by the powerful to keep a check on the people, seems to come down to how power is balanced when lie detectors as invented.
I’m skeptical that it would be used by the people to keep a check on the powerful though. Suppose a lie detector comes out today, and we actually get to use it on Trump and ask him a bunch of questions. How many people will actually change their minds about him, as opposed to make something up about lie detectors so that they wouldn’t have to go through the uncomfortable exercise of changing their minds?
In order for lie detectors to be useful for the people, we also need them to have good epistemics. However, I think public epistemics have been just getting worse and worse in the last ten years. Overall. I feel pretty pessimistic about them being a good thing.