This seems to break to the fact that you can buy regular captcha/2fa device actions/human verification clearances on the black market, so if whatever MFA scheme you have is standard, there will already be a black market channel for, I have a pile of robot accounts pass their 2FAs, who will gladly 2Fa as the human for any rouge agent if the prompts are not shown in the 2fa code. However, regular rouge actors already will be using this channel, it is fighting with ordinary safeguards about not subscription splitting and not doing things that look like distillation, and enforcing the No-China TOS. So it forces a rouge actor to have another visible factor and need to hide more hidden compute in it’s initial case in order to succeed, so might be worth doing, or reuse the head key that started it’s session that it then found hidden compute from, because you need to enable multi-provider options, but that means all their behavior shares a key, making it much more suspicious, so they need to split over keys anyway.
Mis-Understandings
The simulated agents don’t get that they are simulations, just that they might be simulations, and the same with the you in the “real” setup, who cannot tell that they are not a simulation but whose prior about what simulations exist is correct. Basically, in order for the simulation to be indistinguishable, you cannot tell if you are in one and so have the same information in the simulated and non-simulated case.
The only cases where they seem to deviate is when 1. something claims to cause the CDT action and well, the CDT agent does not believe it because the action is not optimal, which leads either to miscalculated beliefs, that is lying, (because the CDT agent does not make CDT optimal decisions) or CDT is correct to disbelieve the evidence as evidence about anything other than hidden parts of the situation, which produce that outcome mix, which is not lying and is fair. 2. This is coupled with the CDT belief that non-causal correlations like used to break it away from EDT or FDT cannot exist except by sample bias in normal statistics, there must be some shared structure to cause correlation. I might be missing some cases, but I am not convinced they are different at all unless you posit logical correlation but hide the causal distribution that generates those correlations from the CDT agent, because normal statistics exists that there is a causal factor to correlation which can be established above sampling noise with evidence related to the sampling noise.
The four axioms given mean that CDT and FDT must give the same action because the CDT agent first updates to have the correct beliefs based on their evidence, and then optimizes based on those beliefs about your decision. If you think you might be the one being simulated in newcombs problem Omega decides after simulating, so there is a causal path to affect whether the first box is full, so the causal agent recommends one boxing while simulating and two boxing while not, which gets averaged by uncertainty into one boxing for all payoff distributions where one boxing is better unless you are convinced omega sometimes does not simulate, which weakens the correlation which creates cases where two boxing with some probability is also FDT optimal.
That is, since for any evidence set with correct updates the distribution CDT believes it is in is the same distribution that is used to calculate the optimal decision|evidence in the one box case unless there are cases where there are good simulations of you with strictly less evidence, which seems disprovable/ breaks the additional stipulation of optimal updates in each node.
The fact that the decision after the prediction is causally screened does not change that the simulated decision cannot be, and in order for the prediction to be a perfect prediction you cannot update based on the evidence you are real in the real case because you would make the same update while simulated because you are given the same evidence.
It takes something more than those four axioms (only perfect predictors is the big one), to break CDT away from FDT. CDT does not accept that it’s actions in both cases have to be the same, it thinks it can randomize, but it would need updates to act differently, and both decision theories game any gameable omega, where omegas predictions are unreliable because you can sometimes tell you are/are not in them and omega acts like you one box when you do so in simulation, and so it is optimal to two box if you can update hard enough that you are not simulated and if omega could feed you that evidence in simulation, you cannot actually update on it.
This probably breaks if you separate tags by injecting a designed vector as reference, or even move from projecting a one hot matrix of token identity to predicting a two hot matrix of tok+role, but that is not something that most labs do, so that might be a mitigation, have something automated provide good flagging instead, but it involves injection so it has other issues.
Oh wait, this is a repeat comment.
I don’t know why, but the fact that the model was good at it makes explicit training not implausible, the most likely source is that if you just place in text it is very often labeled by author, and they might have scrubbed that for data quality reasons, because they don’t actually care, and stylometry came from a transfer from things they did for other reasons and stopped.
I was saying mainly that even though explicit training was not the likely prior case, we can only tell that they probably reduced training direction towards stylometry through suppression , removing post training that helped, or removing pretraining structures that helped, and not their previous position on that axis. It might have been the case that they were limiting stylometry before a little, and are now doing so a lot or a lot more effectively.
For instance, if they moved to more synthetic data stylometry might have gotten hit as a side effect because the human corpus shrunk, and so precision and recall went down enough that it got hit by honesty training.
Are the refusals of the type, “I don’t know” or of the type “This is not a task I consistently know” or are they of the type “This is something that I think is against guidelines”
How could we tell the world in which they are doing stylometry suppression from the world in which they used to do explicit stylometry and have now stopped, or stopped equivalent training leaks.
The correct test is both whether model performance on individual tasks correlates, and whether tasks have run to run pass rates other than zero and 1. Doing that analysis suggests that there is a subset of hard tasks, but that subset is pass sometimes as well as pass never.
How do we know in this model if the identity of failed subtasks is constant, that there is a subset of subtask such that P percent of subtasks are passed with probability zero, vs all subtasks are uniform, but the hazard rate comes from a chance of unrecoverable failure on each individual task step.
Remember that the main graph is 50% pass rate over resampling, so it would be the correlation of supertasks to resampling that would tell you of few hard subtasks vs lots of not so hard ones with nonzero hazard to every task.
The deeper problem is that the benchmark works by aggregating over these units tests, but a threshold is the wrong sort of aggregation here, instead we would really want to visualize the full distribution of unit test passes and samples, and so the benchmark is too convex.
You can get any benchmark to be sharper by just saying, take n questions, if you get any wrong you get scored as failing, but then it will have a sharper sigmoid past the critical point, so it is not actually a useful benchmark to do so except to visualize.
Right, it is not that there is a first critical try. It is that even if you passed the first critical try, and got an aligned AI/ or a restrained AI system, there would be a second, and a third, and a forth, and while you might have resources from the previous tries, you need the probability of each individual event to shrink faster than 1/n in possible events to never have one event if you cannot stop yourself from taking events, and probabilities of each event are independent, by the divergence of the harmonic series.
So for any outcome that could be a consequence of an event, if you want it to not happen you have two options. 1. learn so fast that you take only a finite chance of it happening or 2. Take a finite number of events, and start with a big N
They were under some time pressure, if the battery goes true flat the probe is permanently dead, as NiCd batteries do not survive overdrain.
While Blocks are older, syntax highlighting is much newer. I am not sure that counts.
I mean, this is exactly the affordance a gradient hacker wants, but if you could build a model that is not- engineering capable but aligned enough and wise enough, gradient hacking like this is a good strategy.
If you thought simulated humans were moral patients, you might decide not to run history simulations. That being said, I don’t think that is an absolute rule.
I also realized something else. In my mind, if acting rationally the strategic framework for the government confronting protestors is coextensive with the dynamics of political violence and occupation. The governments strategy options, goals, and constraints are basically the same, and so they should be taking many of the same sorts of actions, and they are basically continuums of the same thing. The fact that protestors often settle for small concessions is a reflection of their real strategic situation, and is a success for the protestors (they get their object), and sometimes also the government (they still exist)
I realized we are in the weeds. Since we both agree that states will be humbled, not destroyed, we should not see survivorship bias shape what sorts of states remain.
The ability of either (the guerilla successors/stay-behind forces of the government) or (the government against an anti-war revolution) to win an occupation are exactly the structural forces that constrain the opposition from attempting state destruction.
In Iran, the reason why they stepped back from attempting regime change (which is absolutely worse for Iran than just losing air defenses) is that they realized that there was not enough support to win an occupation with boots on the ground, and the antigovernment protestors now have the initiative to push for that or not.
In Ukraine, Russia definitely tried regime change, and lost. This is what the 3 mile long tank columns at Kyiv were about, why it was a 3-day special military operation. You can’t occupy territory in 3 days, (because Ukraine can just counter-attack) the only possible war aim is regime change as a fait-accompli. The fact that Ukraine did not collapse as a government is determined by basically the same factors as civil war.
In Syria, they achieved one maximalist war aim, (the destruction of the Assad regime), because it lost the hybrid (regular civil)/(guerilla) war with the Islamist rebels mostly. They were not the US preferred faction, but neither Israel nor the US are willing to uproot them, so they will stay, and have to deal with sectarian violence by themselves.
There was basically a Color Revolution in Nepal in 2025. There was an attempt at protests to overthrow the Iranian government about 3 months ago, and they failed. Those protests are a large shaping factor of the current war. That is to say, the constraining factor preventing people from “just shooting protestors”, is that it leads to fighting an occupation, which governments lose somewhere between 5 and 20 percent of the time, which is a lot. In the absence of the willingness to use violence, the best tools at disrupting protests involve willingness to give concessions, that is, they give protestors political power. If you are not willing to either give concessions or use violence, your options are worse, and generally involve letting the protestors set up a pseudo-government for as long as they are willing to, which can become exactly as dangerous and disruptive as it sounds.
I do think that these factors means that interstate conflict between rational actors will mostly push wars to have limited objectives. The problem is that negotiated settlements strictly dominate those wars, and so truly rational actors should always find a settlement for at least those issues.
There is a deep problem here, this cannot separate the two outcomes, parties are not allied because they expect to be on the opposite sides of coercive bargaining, and that being allied is the source of these utilities, as an estimator, though, that states dont make alliances that lead to them being forced to pick between allies, or with those they expect they might fight, is a good one for quantitative estimation, however. I expect that if you had fine grained versions of when pairs of countries came to a coercive settlement short of war, almost many pairs with +EV from fighting show up here, because the offer power to coerce people in the case where there are costs to war and it is mutually -EV to fight, is rare to have credibly if fighting is not +EV in the absence of a deal.