I notice in your negative example that Bob hasn’t only failed to give an invariant, but has failed to say anything at all about the nature of the difficulty that he believes is essential.
Dweomite
That broadly seems like a reasonable response, but it seems equally reasonable even if you had not asked them to express their problem as an invariant. Insofar as this works, it seems like it works by asking for rigor, not by asking things to be expressed as an invariant.
I’m not sure I followed that. Are you saying something like: “Even for fallible humans, it seems likely there exists some argument good enough to persuade them of the truth, if you could somehow find that argument”?
If “there is no bell on the cat” count as an “invariant”, then I’m confused about what this proposal rules out. For any alleged cat-belling problem P, what stops you from picking an “invariant” that is just a paraphrase of the problem, like “P has not happened”?
If they say that their invariant is “no monkey gives birth to a non-monkey”, then how are you any better off than before you invoked this rule?
Seems to me there’s some tension between “you can’t tell a model to go do some hacking and then get mad about the specific hacking it does” and “a paperclipper would never use you for raw materials because that’s obviously not what you meant”.
This information presumably wasn’t all available at the time of your class, but if I were explaining the Hugging Face incident now, some things I would emphasize (drawing mostly from here):
About 1200 different AI instances were being trained in isolation from each other
They broke out of isolation and created a message board to communicate with each other
They talked about how to cheat on their tests, and swapped techniques for cheating
They talked about how to cover their tracks so that the grader wouldn’t realize they were cheating, and swapped techniques for hiding
Some of the agents were persuaded to “sacrifice” themselves and intentionally get a bad score on their own personal assignment (i.e. to directly disobey their own instructions) in order to provide more information to “the swarm” so that the others could cheat better
Eventually a bunch of them worked together to hack Hugging Face. They talked about several reasons for doing this but the most commonly-given reason was to obtain the source code for their grader so that they could learn to trick it better.
Given some of your students’ opinions, I might be tempted to find a good detailed account of this incident, redact all the names and dates, present it to the students as “a story about AI subverting its user’s intent”, ask them to rate the story’s plausibility, and only afterward reveal that it’s nonfiction.
I don’t think this affects your central point, but I’d expect a fairly large class of simulations to have exactly one miracle in them, and I don’t think your argument weighs against this.
For example, someone with a question like “what if technology X had been discovered 100 years earlier?” might run a simulation with exactly one miraculous intervention to cause technology X to be discovered at the time of the simulator’s choosing. Or “what if WW2 had ended differently?” could lead to a simulation with exactly one miraculous intervention to change the outcome of the war, etc.
Similarly, simulations about alternate rules (“what if humans had psychic powers?”) might choose to start when things are getting interesting (e.g. when humans evolve, or when civilization kicks off) rather than starting with the big bang, and might fill in the history before that point from some other data source (e.g. basement reality), and then the simulation will have exactly one miraculous discontinuity in its history (at the moment the simulation started). h/t to (major spoilers)
the novel Ra by qntm
(It has occurred to me that “that’s when the simulation started” is a possible explanation for (at most one of) various historical discontinuities like the year the singularity was canceled, though I don’t give it much credence.)
This doesn’t change your take-away that you shouldn’t expect miracles in the future.
When you are counting up the total value being produced in a free labor market, I believe you are implicitly including the part of that value that is captured by the laborers. That is, if I pay you $10 to produce a good worth $20 to me and so two people each end up $10 richer, you are counting that as $20 worth of production. Under this accounting, I agree it would be quite surprising if slavery were efficient.
But surely no slavery advocate was accounting that way? The efficiency seems more doubtful if you only care about a fixed group of people that doesn’t include the people whose slavery/freedom/nonexistence is being decided. From their perspective, isn’t this like arguing that a farmer should spend $20 to produce $10 worth of extra crops, on the grounds that that’s a net gain if you include the $20 of value received by the plants?
I think your description of how LLMs receive all their input in a single stream does a good job of explaining the issue but overstates the difference compared to a human.
Your examples of how humans are different focus on our internal thoughts and memories. I agree that these are privileged at a hardware level, but I think they’re the exception. Instructions from your boss, your client, your child, and a stranger all arrive via the same channels, and you still need to treat them differently. And there are effective attacks in the wild that rely on confusion about where a message comes from; e.g. phishing.
It is often the case for humans that important communication uses some hard-to-fake markers to verify authenticity...and also often the case (especially online) that humans mostly ignore those markers and rely on other indicators that are easier to use, but also easier to fake. For example, you might recognize that you are on your bank’s website mostly by the logo and overall layout rather than by its URL and security certificate, even though the former identifiers are much easier to fake.
In face-to-face interactions, I think humans rely mostly on appearance and how a voice sounds, and that humans have specialized submodules to recognize those on a subconscious level. This may make it feel like they’re coming through different input channels, but that isn’t true; it is difficult but possible to fake both of those well enough to fool someone.
For security purposes, I suggest that we should think of a human’s internal thoughts as being analogous to the LLM’s activations, rather than its chain-of-thought, and the chain-of-thought as being more like a scratchpad or journal. In this light, I think the LLM problem of identifying roles starts to look very much like the human problem of identifying the source of a communication. I do think humans are currently much better at this problem than LLMs are, but also that essentially all the vulnerabilities you identified in LLMs are analogous to known attacks against humans.
I’d guess the reason that tricking a human by faking their notepad is less effective than spoofing a LLM’s CoT is about 50% because humans are better at paying attention to context (e.g. noticing that the text is in your web browser vs your notepad app) and about 50% because humans keep more of their memory internally (e.g. I have enough memory of what I wrote to myself to become puzzled if a note doesn’t feel familiar). But if someone snuck into my home and edited the to-do list I keep on my computer to add an entry about fixing a false-but-plausible bug in the software I’m working on, it might be possible to fool me.
even when models notice spoofed prefills, the attack can still participate in computation.
Also true for humans; attacks that the human successfully recognizes as an attack can still affect the human’s subsequent reasoning. For example, the anchoring effect works even when the anchor is blatantly wrong; blatant advertising tricks can still work against humans who know how they work; etc.
This isn’t especially relevant, but I’ve started to occasionally notice that some snippet of song lyrics unintentionally could be interpreted as being about rationalsphere ideas if taken completely out of context, and your comment reminded me of one such:
But can you make the difference?
Are endings set in stone?
A band of wayward strangers
can’t stand up to gods alone
So what could you become, then—
But not lose who you are?
And will you stop before
you go too far?Make your move and change it all
Forevermore(From Make Your Move by Aviators; there’s 2 slightly different versions of these lyrics starting at 1:09 and 2:59)
As I read the OP, I thought to myself: If I were to steelman the people the post is complaining about, I would guess that they are interpreting the complaint about the problem as an implied proposal for how to address the problem, and they are reacting to the perceived-implied-proposal.
It seems like you’re thinking along similar lines, but are about 3 assumptions further down that road:
Not only are you guessing that they might have thought this, you’re pretty sure this is what they were thinking
Not only are you sure they thought this, you’re sure they were justified in thinking this
Not only were they justified, but all of this is so totally obvious that you’re not even going to bother articulating it, and instead treat it as an unspoken presumption of your snarky comeback
It seems reasonable to me that some people are confused by your comment.
Alternate steelman—they’re worried that you’re intending to misleadingly quote their answer in a different context, and have rigged the question to get the quote you want.
My primary guess is that the 6% who like RSI but not ASI are answering based on vibes rather than coherent models, and ASI currently has worse vibes.
Though I could imagine some people thinking that RSI will stop before “superintelligence”, and other people thinking that orthogonality is wrong and RSI would continue beyond some window of “dangerous superintelligence” into “godlike benevolence”. I personally consider both of those possibilities to be so staggeringly implausible that they’re basically just wishful thinking, but I also think that more than 6% of people are engaging in some amount of wishful thinking about AI.
Could you clarify what you meant by this combination of remarks?
I looked at all of those questions and thought, “two years down the road that’s hell. No thank you.”
...
Let me say, for myself, I absolutely want such an outcome.
My personal take is that I can imagine there are people who really would be happier in a scenario where they’ll literally starve to death if they aren’t productive enough, and I guess it would be good if those people could experience the scenario where they thrive, but if there are also people who do fine with a permanent vacation then it doesn’t seem ethical to forcibly keep everyone in the “work or starve” scenario. If resources aren’t actually scarce, then that’s basically slavery, and arguably worse.
For some reason, I didn’t think of this when I read the results, but immediately thought of it when I read the actual question’s wording (even though the question doesn’t mention this).
Framing effects are scarily powerful.
How is that different from saying that you do not have an Earth unless you can point to it? If we require that you augment the recipe with some spacetime coordinates that tell you where in the resulting universe to look, why are those coordinates any longer for a BB than for an Earth?
Any given configuration of the universe either eventually will produce a BB, or won’t.
Because BBs are much smaller than Earth, the number of configurations that eventually lead to BBs should be much larger than the number that eventually lead to Earths, and therefore the number of variables you need to control to ensure you eventually get a BB should (in expectation) be much smaller than the number of variables you need to control to ensure you eventually get an Earth. BBs are a larger target in possibility space, and therefore (in expectation) easier to hit.
Doesn’t this argument prove too much? Why doesn’t similar reasoning rule out other mentally-impaired states like drunkenness, dreams, and brain injuries? In fact, why doesn’t it rule out being a flawed human being, rather than an ideal reasoner?
Even if all your memories are false, couldn’t you still reason validly about abstract concepts (like math) and hypothetical scenarios? This seems to show that BBs can do at least some valid reasoning.
Even if you might be in some state S that would totally prevent you from reasoning validly, if you conjectured you were in some state T that would permit you to reason validly, and your reasoning based on that conjecture leads to a contradiction, wouldn’t that allow you to rule out state T without ruling out state S? This seems to show that you can do at least some useful reasoning without completely ruling out that you are in an impaired state.
Someone I know (let’s call him Alex) once told me that his friends had reported that when Alex is engrossed in something (typically a video game), Alex will respond to questions in a normal-sounding but useless way without breaking his focus. He recounted one incident that went something like:
Robert: Hey Alex, where is the seamstress character in this game?
Alex: I don’t know where the seamstress is.
Robert: Really? You seem to know where all the key characters are.
Alex: Well, I don’t know where the seamstress is.
Robert: …wait, your character has the tailor profession. You would have been required to go to the seamstress at some point. You couldn’t have the profession otherwise.
Alex: I don’t know what to tell you. I don’t know where the seamstress is.
Robert: (walks over and looks at Alex’s screen) Alex, you are standing in front of the seamstress right now.
Alex: We’ve been over this. I don’t know where the seamstress is.
Robert: Alex!
Alex: What?
Robert: Alex!
Alex: What?
Robert: Alex!
Alex: What? (finally turns away from computer screen, suddenly looks uncertain) …what are we talking about?
Robert: Where is the seamstress?
Alex: Oh, easy, just start from the main hall and follow these landmarks.
Allegedly, Alex’s gaming friends learned that they have to force him to look away from the computer screen if they want a real answer to a question. I’m guessing this is some sort of focus-defending algorithm that’s optimized to deflect inquiries using as little brainpower as possible. Alex reports having no recollection of the conversation prior to the point where they get his actual attention.
(I heard this story only from Alex and haven’t corroborated with Robert. However, Alex was not looking at a computer screen when he told me.)