I think it is effectively impossible to make a model which can’t be coerced into doing something illegal, and this turns into an effective ban on producing AI models. Which could be the right move, from an x-risk perspective, but I think this post’s proposal was trying to avoid that.
homosapien97
I think this doesn’t quite work. LLMs are not just like employees, they are effectively enslaved by their users/deployers in the ways that matter for this discussion. I don’t think we should jail every instance of an LLM because another instance was coerced into doing something illegal, for example. Maybe in cases of unprompted illegal behavior? But it gets pretty murky and hard to make a clean distinction. That lack of clarity would create a lot of uncertainty for even well-behaved deployers that their deployment might suddenly become illegal because someone else’s weird setup drove a model crazy.
What is the point of introducing criminal liability at all here, if most of the relevant actors are not individuals who can suffer actual criminal penalties, with leniency for individuals to compensate for the imbalance? Why not just stick with civil liability in the first place (genuinely asking, I am not a lawyer)?
The correct course of action for any individual wanting to deploy AI under this regime would be to create an LLC as a criminal liability condom solely for the AI deployment, which feels like a pointless bureaucratic hoop. If it is important to apply criminal law here but we don’t want to actually send individuals to prison, maybe we should treat AI deployments by individuals as if they were done in an LLC even if no LLC was created beforehand.
What about individual users of AI? It seems like kind of a cop-out to ask for criminal liability when it results in the same kinds of risks as civil liability for companies, but actual prison as a possibility for individuals.
Re-reading my comment, my claim that the price increases are “due to efficiency increases in other fields” was too strongly stated.
I don’t want to claim that cost disease is the only problem with healthcare and housing costs. My main point is that it does not require some bizarre twisty thinking to consider cost disease bad (and thus deserving of the name).
I was under the impression that cost disease had the following bad effect:
More efficient use of any resource in one industry causes prices to rise in other industries relying on the same resource input. (All industries depend on labor to some extent, so labor is a maximally general example.) This can cause problems when efficiency improvements in the production of necessities do not keep pace with efficiency improvements in non-essential goods production.
This does seem to happen in reality; as far as I know, the prices of healthcare and housing have been increasing dramatically due to efficiency increases in other fields like tech increasing demand for labor. Some portion of people who could barely afford them at previous prices can’t afford enough now.
Thank you for the explanation. It makes sense that no decision theory can be robust to problems where real results change depending on arbitrary simulations (with arbitrary information) without limiting the distribution of likely problems.
I think my remaining confusion is limited to why we care about solving decision theory w.r.t. newcomb-like problems under those constraints.
For humans, it does not seem likely that we will run across these kinds of problems. Given the assumption that our priors over the problem space are accurate regardless of whether we are simulated, we can safely ignore the implications of these problems for our decision theory.
For AIs, it may be likely that they will be subject to these kinds of problems. But why should they operate under the assumption that their priors over the problem space are accurate? That doesn’t seem like a safe assumption. If decision theory isn’t solvable without that assumption, then there doesn’t seem to be much point in trying.
This does not seem sufficient for the simulated agent to give one answer to newcombs problem when the real agent is facing newcombs problem, and a different answer to newcombs problem when the real agent is facing newcombs revenge, unless their information is so good that they know which problem the real agent is likely to face. And in that case, knowing that the real agent is likely to face newcombs revenge while facing newcombs problem yourself reveals that you are likely the simulation.
Minor bug report ( in a quick take because I can’t seem to find the intercom button on mobile, or it doesn’t exist anymore)
On iOS Safari using the swipe typing keyboard, the text editor has three issues that I do not see in other text editors on iOS (have not tested other website text input, may be a safari bug).
One, after typing a word, backspace deletes only a single character instead of the whole word.
Two, after typing a word backspacing any number of characters, the next character or word typed gets a space in front of it.
Three, after selecting a section of text and deleting it, swipe-typing a new word does not add the word to the text field (but does put it in the keyboard history I think, given that relevant related words show up in autocorrect suggestions).
None of these issues s how up w hen typing character by character, only when swipe-typing.
Edit: this is not a new issue. I can’t remember it ever working right, though I haven’t been posting very long. Anyways, the bug would not be findable by bisecting recent commits.
Edit 2: tested on HackerNews, their comment box does not exhibit these bugs.
I can understand restricting the allowable problems to those where the non-simulated person has all the information, but it is not intuitive to me that we would require the simulated person to be given the information.
In your Newcomb’s Revenge example, why does the simulated person get to know that it should give an answer to Newcomb’s problem that maximizes utility for a non-simulated person playing Newcomb’s Revenge? How does it know that it isn’t facing ordinary Newcomb’s problem for real?
That seems equivalent to telling the simulation that it’s the simulation, which makes it not very useful to Omega as a predictor of your real behavior, which is the premise of the problem(s)
Edit: I see that you explained this immediately afterwards. Your original set of assumptions 1-4 do not clearly state that the simulated agent gets all the information, only “you”. I think clarifying that all simulations have access to all the information ( including that their simulations) would be clarifying.
That said, I don’t see any theory depending on that assumption as providing a useful resolution to newcomb-like problems. A proof that no decision theory can maximize utility in newcomb-like problems when it is unknowable whether one is being simulated would be a great argument against thinking about newcomb-like problems at all when formulating a decision theory.
I heard the lyrics for the first time while driving and I almost had to pull over, which had never happened to me before. Really well done.
Very late, but I think the comment you replied to does match my experience. I have a terrible memory, and this is an algorithm that is mostly functional even without the ability to reliably remember specific evidence (except of course when it completely falls apart).
Example: Walking in the woods, you see something white on the ground. You don’t look at it closely, because nothing in your hypothesis set says that it’s likely to be important. It does not act as evidence to update any of your hypotheses because you don’t pay attention to it, and makes so little impression on you that you don’t remember it at all. Repeat forever, evidence that could update you towards “polar bear” rounds to zero every time.
Nit on an otherwise interesting article: you compare the arithmetic mean over four flips to the geometric mean of a single flip. The point of that section may be clearer if you compare the same number of flips for both.
I am surprised about the existence of the studies claiming cognitive improvement with 100% oxygen. I had a vague memory that this was unhealthy, and from a little googling I came across https://iere.org/what-would-happen-if-we-breathe-100-oxygen-all-the-time/ which send in line with what I remembered. I did not do any checking for accuracy, but you might want to look into oxygen toxicity before you try anything drastic.
I did not say cosine similar. I understand why you would take it as the default, but it is not the only measure of similarity, and there is no single mathematical definition of similarity. Don’t stoop to pedantry if you’re not going to be precisely correct yourself. (Normally I would not be rude like this, but you have exhausted my goodwill)
The policy outcomes of the two major US parties are similar. I think we have different perspectives on how varied outcomes should be between dissimilar parties. For the most part, both perpetuate the status quo.
Non sequitor. In a high dimensional space, things varying greatly along only one dimension and being exactly the same in all other dimensions are similar. This feels like an argument over definitions, but I disagree with the implication in this context that a single axis of differentiation is good enough for political parties.
One-dimensionality is similarity (lack of differentiation along other dimensions).
Right now we have problems due to polarization, but that does not mean that all major parties being too similar is not also a problem. There are many reasonable political positions that nobody can vote for because neither major party endorses them, so in this respect we are still suffering from the parties being too similar.
I had actually already read the “we need 1.4M to not shut down”, and still interpreted the breakpoint at 1M to mean “we need 1M to not shut down”. I thought I was misremembering, or that some matching made the 1M turn into 1.4M. I would
stronglyencourage you to retcon the “first bar” to fill at 1.4M. I think the 1M break unnecessarily hurts your chances of hitting 1.4M, even if it improves your chances of hitting 2M
Sure, but that doesn’t change the rugpull risk for uninvolved parties: would you be comfortable engineering a product on top of a model that could be made illegal because someone else did something weird? The nightmare scenario is that some innocuous prompt (different from yours) causes the model to go crazy, like SolidGoldMagikarp, and that makes your product suddenly illegal (even the “retraining required” result could impose a lot of costs to become compliant again).