frmsaul
Pretraining an LLM without mentions of consciousness
Misaligned AI in the Bronze Age
The singularity began in 2026, when Claude wrote its own harness.
The singularity began in 2022, when the first ChatGPT finished training.
The singularity began in 1977, when silicon moved into the living room and taught kids its language.
The singularity began in 1945, when von Neumann switched on ENIAC and computation went electronic.
The singularity began in 1602, when a group of Dutch merchants discovered that capital could be incorporated and set loose to optimize for its own growth.
The singularity began in 3200 BC, when the people of Uruk invented writing and immediately began producing pre-training data.
The singularity began in 60,000 BC, when a band of hunter-gatherers invented recursive grammar and cultural evolution was born.
The singularity never began, and it will never stop.
Markets that are more “fun” might have more trades and be less predictable, so they have higher volume and higher brier.
I‘d be curious if you still see this pattern if you control for “fun”.
yeah agreed. That’s a better term to use.
This is a good perspective. I agree.
Thank you.
Re: shaky reward hacking. Yeah I agree. There are a lot of things going on here and it’s unclear to what extant reward hacking plays a role in the pavilion decisions.
My model is basically this:
reward: Being a great country with high quality of life for it’s ruling elites
proxy: Being perceived as a great country by it’s citizens and others
hack: spending way too much money on symbolic things rather than improving infrastructure, trade, human capital, etc.
The reward and the proxy are obviously correlated. The relationship is also causal, e.g being perceived as great by it’s citizens make the country more stable. So it’s arguably very rational for the ruling elites to invest a lot in prestige.
This is a good framing.
> both authoritarian societies and democratic societies seem incentivized to over-emphasize the relative merits and successes of authoritarian societies over democratic ones
Why do democratic societies have an incentivize to over-emphasize the successes of authoritarian societies? Is it a function of the “opposition” trying to win elections? (“The Prussians are beating us! This is because the current government sucks, elect and we will prevail.”)
I think I agree. Re: the space race. There is a little bit of a “duel use” element to it. The same technology that takes a satellite to orbit can also bring stuff from the US to Moscow really really fast.
Reward Hacking at the 1937 World’s Fair
I really enjoyed the post. In particular, I loved this part below:
```
In some sense, reward hacking coming from RL shouldn’t be too surprising, as RL trains models to take actions that get high reward, and reward hacking is sometimes a good strategy for getting high reward. More concretely, RL may teach models certain behavioral traits that make the models more prone to reward hacks. RL training can teach models to be more persistent, to find creative solutions, to try solutions that are unlikely to work if there is no better alternative, and potentially to even think about how they are evaluated, as these traits are useful towards accomplishing tasks in a diversity of environments.
```
It made something[1] really click for me.- ^
Not sure what to call it exactly, maybe: “generalization of reward hacking” / “Reard Seeker Selection”
- ^
> We have no counterexamples or causal experiments.
That is the nature of politics / economics / culture. There are no ways to conduct meaningful experiments on policies, but we still need to make policy decisions, so we need to have other ways to build causal models of the world. There’s a bunch of literature around how to do that. Daron Acemoglu won the 2024 nobel prize in economics for innovating some of these methods (specifically in the context of “institutional selection”)
> the forager → farmer → industrial → modern “progression” has included some return to forager values, and some brand-new directions that probably just wouldn’t work without the level of wealth and automation we’ve managed.
This is a really interesting point, can you expand on it more?
Both the Meiji restoration and the Alexander II reforms greatly increased the liberties of the average person. The fact that they made the state more powerful means that “liberalization” is adaptive and only states that do it “win”.
I agree that the “rise of China” is an interesting counter-point. Xiaoping’s China adopted some liberal ideas (free markets, some local democracy, some free-speech, some rule of law) but didn’t adopt “the entire liberal package”.
I imagine that de-centralisation and capabilities slowdown can go together. The government can heavily regulate the labs and limit capability growth, and the labs can compete on other things e.g costs, speed, “values”. What do you think?
Is Progress Inevitable?
frmsaul’s Shortform
I Haven’t Thought About the Blood Pipeline in Years!
As a student at the Technion, engineering exams evaluated both my capabilities and my ethics. What does it mean? It’s all downstream of this story:In 1961, the late Professor Haim Hanani was appointed as the vice president of the Technion. Once appointed, Hanani proposed to start teaching also humanities. The other professors did not agree. They claimed that there’s barely enough time to teach “hard” sciences.
Hanani wanted to prove his colleagues wrong. He gathered a hundred students, and gave them an exam with one question: what technical information do you need to plan a pipeline to transport blood from Ashdod to Eilat? The two cities are located 250km apart.
The students started working on a solution right away. Using drawing boards and slide rules, they suggested to measure the topographical situation along the route, check the pipe layout, and test the corrosion resistance.
When they finished and submitted the test, Professor Hanani announced that they all failed: “I did not ask to test your ability to plan a blood pipeline, but to examine your moral sensitivity. None of you asked whose blood will flow through the pipes, or who is asking to build it in the first place”.
Nowadays, Technion professors will sometimes hide ethically loaded questions in exams, my responsibility as a student was to find them and point them out instead of “just following orders” and answering them verbatim. So every time I faced a question on a problem-set or an exam, I spent a few seconds trying to decide if it’s a “blood pipeline” question or not. “Does this linked list represent a blood pipeline? Could this linear program be used to optimize the intake in a concentration camp?”. I don’t think I was ever actually presented with such a problem[1], but the blood pipeline was always in the back of my head.
In the language of ai alignment, you can say I was evaluation aware, maybe even evaluation paranoid. After leaving school, this habit didn’t really stick. I haven’t thought about the blood pipeline in years.
- ^
If I was, I didn’t notice
- ^
I like your theory. It would be interesting to see some mechanistic interpretability studies of this phenomena.
We Need to Get Serious about Uplift Studies
Yeah i totally agree, this is probably the right approach. I recently wrote about this type of benching.
Wow this is awesome! Yeah we would love to talk. Can also grab coffee or something (assuming you are in the bay)
(my contact info is on my website frmsaul.com)