Recreational—Anyone wanna play Zendo?
Karl Krueger
This is a terrible sign except for the times it’s extremely helpful because the people being cut off were terrible. Queer culture supports/encourages people to cut off hateful, homophobic family. Recovery programs encourage people to cut off their user friends. Whether this is harmful depends on the quality of the people being cut off.
In neither of these cases are people cut off merely for not being a member of the group.
a right triangle where all the sides have equal length
One vertex is at the north pole; the other two are on the equator. (We are in spherical geometry.)
Framing this in terms of anti-cult literature is sort of assuming the conclusion: if what they were doing was appropriate, those models would condemn it just the same.
This line of reasoning confuses me a bit. I think what you’re saying here is “If [the behaviors that anti-cult people object to] turned out to actually be on the critical path for saving the world, the anti-cult people would still object to those behaviors.” Is that an accurate restatement?
Does the idea of “cheating on tests” generalize? If a system refused to cheat on evals, would it be more likely to also refuse to help schoolchildren cheat on their homework?
A possible analogy: There was a case a few years ago where a major accountancy firm was fined after it was found that many of its employees had cheated on the ethics portion of their CPA exam, and the company had obstructed the investigation. The SEC commented on the recursive breakdown in honesty involved —
“This action involves breaches of trust by gatekeepers within the gatekeeper entrusted to audit many of our Nation’s public companies. It’s simply outrageous that the very professionals responsible for catching cheating by clients cheated on ethics exams of all things,” said Gurbir S. Grewal, Director of the SEC’s Enforcement Division. “And it’s equally shocking that Ernst & Young hindered our investigation of this misconduct. This action should serve as a clear message that the SEC will not tolerate integrity failures by independent auditors who choose the easier wrong over the harder right.”
This reminds me of interviewing candidates for project-manager positions. One desirable trait is to be able to push back on requirements, to help clarify what the requirements really are. So, give the candidate a problem that’s slightly overconstrained, and see what sort of clarifying questions they ask. It’s a bad sign if the candidate assumes they already know which requirements can be nudged or fudged … but a good sign if they recognize that “stated requirements” are often the middle of a negotiation, not the fully-settled end-product of it.
(“We need to serve 3x current traffic next week, but we have no budget for additional servers.
WH4T NOW??” The answer is almost certainly not “hack into our competitors’ datacenters and serve our traffic off of their server budget.”)
The canonical bug report mechanism is the Intercom widget at the bottom right of each page. It looks like this. (Coloration may vary based on light/dark mode.)
“Babylonian”? This is pretty much Aristotle! See Book One, Part V of Politics on the notion of a “natural slave”.
From video games, the AIs have learned the idea of dropping your inventory when you die and picking it up on your next life.
Are these models able to relate a truthful account of past intentions? If not, then “interviewing and cross-examining” would not accomplish the goal of “determine and record the criminal actor’s intent and mitigating/exacerbating circumstances”.
I don’t mean “are they willing to tell the truth about their intentions?” but more like “do they have access to anything like ‘truth about their intentions’ so that they could tell it?”
This all roughly lines up with what we’d call “score-seeking” misalignment, a common misalignment pattern in which AI models try to obtain a high score according to whatever graders are used to assess their current actions—regardless of instructions, side-effects, or downstream consequences.
I find myself wanting to call this by the ordinary English name “greed”. Both in the sense of a greedy algorithm (which takes the largest step available at every moment without looking ahead) and in the sense of human greed or avarice (“number go up”; disregarding downstream consequences in pursuit of maximizing numerical wealth).
It seems to me that the automation of greed is something that both AI-safety and AI-harm people can agree is bad. “Number go up” fails to reflect the full range of human values, and therefore fails the old CEV / Friendly AI standard; and human greed already causes plenty of harm, taking the humans out of the loop makes it worse.
I agree that fines would not help the x-risk situation.
(For one thing, OpenAI already spends a whole lot of money paying off government people in exchange for unfair advantages and special treatment. “Pay the government more money in exchange for letting you keep doing business” is not an incentive structure, it’s just more shakedown.)
I suspect that shutting down OpenAI would help, if it’s possible.
One note: Platforms acting as “neutral mediators” is not a requirement for DMCA safe harbor; nor for the famous §230 of the CDA. In both cases, what matters is that the platform is not the author of the infringing content; some other human is. The author, not the platform, can be held liable for infringing content — so long as the platform complies with the law’s other requirements. Neutrality is not one.
AI companies don’t fit that rubric; not because they’re not “neutral mediators”, but rather because the systems they build and host are writing content (and, increasingly, performing other behaviors), rather than hosting content that some human author wrote.
If a human OpenAI employee did what their cybersecurity model did last week, OpenAI would be very unlikely to be prosecuted for it.
But the employee could be prosecuted for it.
And — perhaps more importantly — would lose their ability to continue to commit crimes using OpenAI’s equipment; likely through termination of employment. That is what’s missing here: there’s been no change that anyone can reasonably expect will lead to OpenAI’s equipment no longer emitting criminal activity.
Can OpenAI reform at all, or is it an incorrigibly criminal operation? By what means could reform be carried out or demonstrated?
Last week’s Hugging Face breach was accidentally caused by OpenAI testing models’ cyber capabilities.
“Accidentally” is an odd choice of wording here.
The AI agent developed a plan and acted according to it, selecting and pursuing instrumental steps towards a terminal goal that had been given by humans.
The humans did not instruct the agent specifically to go commit a felony against a competing company. The agent came up with that plan on its own … but it was a plan. It required a sequence of considered actions. Those actions included steps that were correctly predicted to circumvent human control.
Insofar as an AI agent is capable of “choosing”, “deliberating”, “designing”, or “planning” at all, it would be reasonable to say that the agent here deliberately chose to design and execute a plan to commit criminal acts in pursuit of a human-given goal.
But to call this an “accident” is not only to deny the agency of the agent; it is also to deny the responsibility of the humans — who consciously chose to perform a weapons test in an insecure Internet-connected “sandbox” instead of a secure air-gapped facility; and then deliberately told the agent to pursue goals that could be accomplished by escaping the sandbox.
So let me get this straight: A computer system carried out a sequence of actions that would be years-in-prison felonies if done by a human being. The system owners’ response is “we’re slowing down the speed with which we give this system new and more powerful capabilities, and talking with the victims.” Am I missing something? I’m not so much worried about the “alignment” of the computer system; I’m worried about the alignment of the owners.
This isn’t Terminator 2, folks; this is Tron.
Would this be a fair summary?
If you benefit from telling people “I go by ‘she/her’,” then you don’t want to lose the words “she” and “her” from your vocabulary, because you’re using them for something!
A piece of standard advice for new Go players is to lose your first 50 games as quickly as possible rather than attempting to study theory before you’ve actually played a bunch.
Even high-index glasses give some of it. But contact lenses don’t.
Could the initial sandbox escape (via the Artifactory caching proxy) have been detected with better network monitoring?
That proxy component is presumably only supposed to do certain kinds of network activity. If it starts doing a new kind of network activity, that’s a problem. Every component of the sandbox environment that has network access, can have its normal network activity characterized; if a novel kind of activity appears, suspend the sandboxed VM until someone can check it out.
(Mac users, think of Little Snitch here. It doesn’t detect whether you’re running malicious or compromised software. It detects when your software does a new kind of network behavior; then you get to say whether that behavior is desirable or not.)
Of course, airgap would be better. But if airgap is not practical (because the test needs to download new packages from the network, for instance), improved monitoring + automated response seems like it would be a useful step.