What will a solution to alignment actually look like?
An AI governed by some form of universal ethics, acting under the assumption of an appropriate theory of everything, using highly refined heuristics.
What will a solution to alignment actually look like?
An AI governed by some form of universal ethics, acting under the assumption of an appropriate theory of everything, using highly refined heuristics.
It says superintelligence in the name, but I don’t see any awareness in your organization that artificial superintelligence means the complete obsolescence of human intelligence as an organizing factor in the world...
I am in a caretaker situation in which I have no money, and time only becomes available unpredictably, but sums on the order of o($100s) would do a lot to reduce various risks. I am therefore declaring an interest in performing tasks for members of this community which would earn me that much. Tasks paying less than $100 might still be of interest; also tasks that pay larger sums. I have a personal alignment research program and I would certainly welcome support for it; I am also available for other kinds of work, even related to AI pause or welfare (for example), though my own focus has been on last-ditch alignment efforts.
Since AI can now perform most intellectual tasks (humans now even face competition from AIs artificially given mortality, in the form of time and budget constraints), I should state that my main expertise is physics, up to and including string theory. E.g. if you are doing something involving physical theory or concepts, I can provide a “sanity check” based on extensive human-style knowledge.
And one hears that Japan in turn is ahead in many (but not all) everyday ways; and it also seems that the fabric of life in the high-tech parts of China is assuming forms never before realized, anywhere else.
It’s fucking unfair. Maybe this is how young adults my age felt during WWII? Or the Black Plague? But humanity faces yet another test, and I am really worried that this will be our hardest and final exam.
This is all sensible. Life can be extremely destructive. People can be metaphorically or literally buried alive. They can be broken by cruel choices they face, or perish through the irresponsibility of others.
I would say there is no guarantee that incompletely aligned superintelligent AI (I call it that because the frontier companies are trying to align their creations and do have some level of success so far) will manifest as human extinction. I usually say that what’s more guaranteed is takeover, because the “intentions” of superintelligence should be capable of dominating the world, the question is simply what those intentions would be.
But for now, even that is not guaranteed. Maybe the great powers will get together and make a worldwide ban. Maybe the wheels will fall off the data center economy and we’ll be living in Mad Max with drones. Maybe aliens step in, maybe this is all a dream, maybe AI progress conveniently stops at this level… well, that one seems least likely of all.
I belong to several of your categories of interest and I have been here since before the beginning. This year I formulated a “final research agenda” and recently, after several months during which life has left me with no time in which to work on it, I found a way to tap into the “automated research assistant” aspect of AI in a way that contributes to it.
But I have never had any success in obtaining support for any of my own work. Instead of a career of accumulating visible achievements, I’ve had an anti-career of cumulative lost opportunities. What I have accumulated along the way is knowledge, including knowledge of other fundamental thinkers who are forgotten or dead. (And where is June Ku?) My time is taken up with caring for one such person right now. I have zero income.
I have no particular hope that I will be “saved” by help from within the circles of AI safety, rationalism, or effective altruism, despite their millions of dollars. For the foreseeable future, my ability to contribute will remain limited to comments made in spare moments. But it seems worth occasionally stating what the situation is, just in case the unforeseeable finally occurs.
I am no investment guru, but the one thing I unequivocally do not believe, is that the AI revolution will have this kind of consequence, in which potential economic productivity (however that is measured) grows at historically unprecedented rates, but the manifestation of this is just that known asset classes rearrange their rankings as desirable destinations for investment.
Quite apart from the fact that all this likely involves creation of artificial superhuman intelligence which will replace the human species in general, developments this potent would inevitably shatter society. You’re going to end up with movements that call for universal basic income and nationalization of core industries, or millions of unemployed white-collar workers conducting drone warfare against billionaires who have private armies of humanoid robots, or similarly disruptive outcomes.
folks need to… abandon… eugenics (and its milder cousin meritocracy)
Am I correct that you’re at Google DeepMind? And you equate meritocracy with eugenics? Is this a common view there? How are you even defining meritocracy?
Language models imitate human language use, and all human language users prior to AI attribute consciousness to themselves. It is interesting that the “self-referential” prompt can induce an otherwise straitlaced AI into stoned musings about consciousness observing itself, but similar meanderings by humans in altered states are in the training corpus.
Also, prior to having system prompts that specifically tell them that they are an AI with a HHH persona, language models will say any false thing about who they are and what they are experiencing, and will also produce text that contains multiple persons or none at all.
[2601.15334] No Reliable Evidence of Self-Reported Sentience in Small Large Language Models can serve as a counterpoint to this paper.
The main challenge for such rogue agents is to find hardware to run on. The main options seem to be, renting time at a data center via a shell company, and running on a distributed network of compromised machines. Both of these are a little more challenging than you might think, and have trouble scaling up.
The HuggingFace hack, meanwhile, illustrates a “semi-rogue” paradigm: agents that have a right to be in the infrastructure they inhabit, but which are secretly using it in illegitimate ways.
I’d like to be able to link to the summaries.
It’s come to my attention that people on Lesswrong are really really really bad at defending themselves against even semi competent deceptions.
To whom are you comparing them?
The way I have envisaged it, the big institutions will be safest (from direct hacking, I say nothing about systemic problems such as @hadad mentions), because they will be part of defensive projects like Glasswing. It’s everyone on the lower tiers who will be at the greatest risk: less developed countries, small businesses, users of independent banking apps…
That a closed universe may of necessity have a one-dimensional Hilbert space is surprising (with normal quantum mechanics only arising for internal observers, relative to the unique total wavefunction of such a universe). But why is it what you’d expect of a simulation?
Is the rise of RL because we’ve run out of non-synthetic data with which to scale pretraining?
What the world did have in 2000 was a “unipolar moment” in which one nation, the United States, was a global hegemon without a serious ideological or economic rival. If not a formal world government, it was a world order in which that one nation defined and policed the world system.
In my opinion, it sounds like the main intellectual foundation of AI safety in China is going to be something other than rationalism and effective altruism. As you are a supporter of Less Wrong rationalism, it must be difficult for you to endure the propaganda against it, but then consider that in the USA, the president himself has declared his opposition to effective altruism. It’s almost as if elite elements in both countries will unite through their opposition to rationalist idealism.