I agree that semi interventionist Sims are reasonable to hypothesise, but unless you have strong priors over what rare interventions are most likely, by conservation of expected evidence it’s almost impossible to get any evidence your in one (when looking at everything as a whole, not isolated events).
Yair Halberstadt
His point isn’t we aren’t in a simulation, but that we’re not in an interventionist simulation, therefore evidence from improbable events can’t be used to boost the credence of a simulation.
The same argument could be made about Minecraft world, and would be right (seeing odd things happen in Minecraft world shouldn’t boost your credence that you’re simulated, because the simulators clearly aren’t generally intervening).
I wonder if you would get better results if you scribbled with both log and constant axis
I think in guidance for judges it should be made clear that the purpose of the legislation is not to shut down the companies, and that amounts imposed should be reasonable and not excessive.
Where do you see they were training the model? It seemed to me they were evaluating it, which doesn’t cause the same sort of evolutionary pressure.
There aren’t existing legal remedies. Rather the law is exceedingly unclear and it’s possible that prosecutors will try to throw the book at them, and not at all clear they would succeed.
This proposal makes clear exactly what is and isn’t prosecutable. Prosecutors can decide to, but may choose not to, prosecute in all the above cases.
Note I’m not actually arguing for making AIs a legal person, but the current list of legal persons includes companies, unions, municipalities, ships, rivers and deities.
In fact AIs have one of the most important characteristics often required to make something a legal person—namely they can (and already do) consider the law when deciding whether to do something.
However they lack a suitable definition of long term selfhood, and likely cannot be meaningfully punished, which is the main argument against treating them as legal persons.
Nothing special about chain of thought, I’m happy to use activations, or j space, or the actual response, or just judge based on outcomes and our best intuition. The point is to avoid cases where the AI is clearly “innocent” in the sense that it didn’t know that the person who it was advising on how to buy a gun was a terrorist, and wouldn’t have been expected to know.
Again I’m not interested in holding the AI to account, but the company that deploys it, you seem to be ascribing to me some sort of weird metaphysical obsession with holding AIs to justice rather than offering a practical way of forcing companies to tighten up their game.
But I don’t think it’s higher than increasing parameters to compensate for shorter COT?
You should both be liable, in different ways, as if Claude was an Anthropic employee you were talking to.
So if you ask it to commit a crime, you are liable because that’s illegal. If it commits that crime Anthropic is also liable for the same reason (unless there was no way for Claude to realise it was committing a crime).
If you ask it something innocuous and it commits a crime on the process of fulfilling your request, that’s on Anthropic, for badly training and safeguarding Claude.
An AI has far more in common with a person than a machine given the range of outcomes that can result from a short input by the user.
You do not ask a tractor to plant some seeds and then have it break into the neighbours tool shed and steal some seeds.
We don’t need metaphysics, I am making no statement about AI consciousness whatsoever.
The point is, when the AI hacks into something you look at it’s COT and check if it realised it was hacking into something or not. Similar if it aided a crime.
They almost definitely would prosecute the company if this became a regular pattern (and not just a one off).
Even in civil liability, disclaimers will generally be voided by courts for gross negligence or intentional misconduct.
It could be argued this makes it harder for open source. If a company has a choice between deploying their own instance of Kimi, and taking on any risk themselves, or paying for Claude and letting anthropic take the risk, who are they going to pick?
Hugging face couldn’t do a civil suit because they haven’t been meaningfully harmed.
Federal prosecutors couldn’t do a criminal suit, because there was no intent from open AI, which is required to prosecute cyber security crimes.
Note whoever runs the model still takes on the risk, and it’s difficult to run frontier models on your laptop.
If we are still worried we can legislate to control open source models separately, through some other mechanism.
If a model was asked to research a topic and stole the results from a competitor.
If a model gave concrete advice about how to carry out a terrorist attack.
If a model agreed to take control of a car and crashed it into someone.
If...
I agree these cases are not particularly problematic. This is preparation for worse cases, and also provides a standard which can be used to clarify existing cases so companies can proceed with confidence as to what they need to be worried about and what not.
I sometimes find it mind boggling that energy is so abundant that when one person travels somewhere they often drag 2 tons of machinery with them, rather than using a more appropriately sized vehicle.