Other examples: A Covid-19 pandemic also seemed intuitively highly implausible in early 2020 despite strong evidence that it could no longer be contained. Similarly with Russia being about to invade Ukraine in late 2021 / early 2022. There should be a term for this: *normal world bias*
cubefox
A counterexample is a variant of the typical mind fallacy: if you perfectly understand something, it feels so obvious that you tend to forget what was not obvious before. So professors are regularly worse at explaining a difficult concept than teaching assistants: because the latter still know what makes it hard to understand.
This Twitter user is known for having an undisclosed method for accessing the CoT of OpenAI models. He sometimes posts snippets. Recent examples.
Well, if RL “really worked”, yes it seems like it should almost certainly invent and reason in an alien language.
What would be the advantage? Inventing new languages at least doesn’t seem to increase our reasoning ability compared to using existing ones. For example, it seems unlikely that people can reason better with any artificial languages like Esperanto or Interlingua (or Python, UML, etc) than with English. I would expect that if it were possible to artificially create a language that is substantially better for thinking than existing languages, we would have already created such a language.
Of course humans can also reason to a significant degree in latent space, without language, but latent reasoning (Neuralese) is different from reasoning in an artificial language (Thinkish).
From hovering I can see there are just three accuracy votes (for a total of −11) so far, so someone apparently used strong downvotes.
I didn’t mean to suggest otherwise. Probably bad wording.
The issues might already be present during SFT (instruction tuning), before any RLHF or RLVR is applied. Base models, being pure token predictors, don’t have a problem with faithfully extrapolating style, but they don’t reason and therefore can’t plan ahead very far, so any complex plots are out of reach.
Damn, you beat me to it by seven years.
I’m impressed that you managed to put a hyperlink inside your LaTeX.
A nice definition of “paradox”.
David Lewis famously observed that one man’s modus ponens is another man’s modus tollens. The first man argues:
While the second man argues:
But Lewis forgot about the third man:
each major government
Particularly the US government. Other governments will consider themselves lucky when they manage to convince the US government to get access to frontier models.
Well, I just used a semantic distinction to resolve the apparent incompatibility between CDT and FDT, so this semantic argument seems to be the opposite of fruitless. Indeed, my point is that not looking properly at the semantics is part of the reason why there is a dispute in the first place.
(Just to be clear, I think your post is great! I was just commenting on this narrow point which I think is important.)
Critics judge the branch; defenders judge the policy; and I don’t think the word “rational” settles which is correct.
It depends on what we apply the word “rational” to. Are we asking whether a fixed policy is rational? Or are we asking whether a decision in a particular situation (“branch”) is rational?
It is no contradiction to say that a policy which never pays ransom is rational and that, in a situation where you have been blackmailed, the decision to pay the ransom is also rational (assuming this doesn’t cause the probably of future blackmail attempts to increase). A rational policy makes irrational decisions in certain situations.
CDT can be viewed as describing which decision is rational in the current branch (decision situation), and FDT can be viewed as describing which policy is overall rational to have across branches. Viewed like this, the two theories are perfectly compatible, they try to answer different questions.
The sentence “A rational policy makes irrational decisions in certain situations” sounds contradictory, but isn’t. This is arguably why Newcomb’s problem feels paradoxical.
A few seconds later Gemini 3.1 Pro just jumps straight in to take over its younger sibling’s computer without asking…
Gemini 2.5 Pro is actually the older sibling, because it was born earlier.
Then the question is: the better “outcome” of what? Of an individual decision situation? Or of generally implementing some decision algorithm? The answer will be different accordingly.
The question is what counts as “better decisions”. If the best decision is the one with the highest “expected utility”, it depends on how this concept is made precise, and since every decision theory does this differently, every decision theory is the “best” according to itself, which is not very interesting.
So the actual question is: what are “better decisions” according to our informal human intuitions? This is sometimes exceedingly hard to determine. Newcomb’s problem is the prime example. It’s called a problem because every paradox is a problem, and Newcomb’s problem is a clear case of a paradox. From Robert Nozick’s original paper:
I should add that I have put this problem to a large number of people, both friends and students in class. To almost everyone it is perfectly clear and obvious what should be done. The difficulty is that these people seem to divide almost evenly on the problem, with large numbers thinking that the opposing half is just being silly.
Given two such compelling opposing arguments, it will not do to rest content with one’s belief that one knows what to do. Nor will it do to just repeat one of the arguments, loudly and slowly. One must also disarm the opposing argument; explain away its force while showing it due respect.
My own take on solving this paradox (I should write a post about this) is that theories like CDT, and theories like UDT or FDT try to answer different questions. The former try to answer which action is the most rational in a specific situation, while the latter try to answer which decision algorithm is the most rational across all possible decision situations. (Though I think “useful” might be a more precise term here than “rational”.)
Cases like Newcomb’s Problem seem paradoxical because they are cases where a rational decision algorithm takes an irrational action. Which sounds like a contradiction, unless you notice that the objects of the predicate “rational” are different: local actions versus a global algorithm. Similar concepts of local and global instrumental rationality have been discussed in the past in contexts like MAD.
There are other reasonable definitions of “fair” according to which Newcomb’s Problem is clearly unfair: the predictor punishes people who pick the most useful option in the given situation, since “usefulness” is a causal term. So it punishes specifically agents who pick actions according to causal expected utility. There are other cases of uncontroversially fair problems, like tragedy of the commons and prisoners dilemma type situations, but Newcomb’s Problem is not one of them.
CDT does not, in general, have a good way to pre-commit to actions. Nor does EDT. Since pre-commiting to actions is extremely common in real life (“I will hire you if and only if I think you won’t slack off and cause trouble for me”)
I disagree. The example is not a case of a pre-commitment. Pre-commitments move decisions from the future into the present, such that your future self executes the decided action with slavish certainty as if under hypnotic suggestion, such that nothing is left to decide in the future. Humans generally can’t do this. We may think we can “pre-commit” to washing the dishes tomorrow, but when tomorrow comes, we still have to decide whether to wash the dishes or not. Past “pre-commitments” are then just recommendations about what to do from our past selves that we may safely ignore, like we can ignore recommendations from other people.
Is it intentional that this isn’t shown on mobile?
An unrelated point: these pictures seem obviously bad. They look like children’s drawings, or drawings of a random amateur. The artist is blind, so this is understandable, but it doesn’t make the end result any better. It seems that the museum is more interested in promoting certain artists who are deemed deserving of attention than in promoting good art.
The whole thing looks like a “The Emperor’s New Clothes” situation: High-brow curators and visitors looking respectfully at the stuff in front of them until some random kid (no doubt dragged in against its will) just says “this looks like shit and also weirdly perverted”.
Speculation: Are we perhaps already seeing the indirect effect of diffusion model art here? When photography was invented, artists pivoted to unrealistic art. Now with text-to-image models getting better and better, even technical drawing prowess for unrealistic art is becoming unimpressive. So artists and creators have to find the interesting aspects of art elsewhere, like in the unusual personality or circumstances of the artist.