I write software for a living and sometimes write on substack: https://taylorgordonlunt.substack.com/
Taylor G. Lunt
In a recent post, I mentioned FrontierMath Open Problems benchmark as one of the few I take seriously for measuring intelligence (as opposed to e.g. coding benchmarks). My mostly-a-joke LaughBench benchmark is also starting to show minor results. This is the first time I’m sensing that the AI models I use have some kind of spark of intelligence inside them.
True for people as well.
The distinction between “decent” jokes and ones that actually make a person laugh seems like an important one to me.
I unfortunately don’t remember what the jokes were.
LaughBench
This seems like a great possible explanation for why LLMs are so reward hacky these days. Using latest models, I feel like I’m constantly being “worked” by the model, and lies/deception are common. I agree the problem is probably something with the RL setup.
I have also considered ways in which AI could learn unaligned strategies even if the strategy is not immediately rewarded. I wrote about some of my thoughts here.
I think it only counts as a warning shot if it’s scary relative to the previous warning shots.
That said, I probably agree. Thank god for slow takeoff.
I make TikTok videos and talk to normies about AI safety, and people seem more interested in this incident than they have been about anything I’ve said in the past. There is an element of “shit just got real” that is inherently convincing.
However, many people either think it’s a PR stunt or that the AI was specifically instructed to hack Hugging Face, and that AI cannot do things unless specifically told to do so. Any attempt to communicate with the general public needs to deal with these misconceptions.
If this is true, maybe you could funnel people toward 1:1 work. I understand you do 1:1 work, but you can’t see everyone yourself. I wouldn’t even know what kind of person to look for. A zen monk? I am clueless. There are millions of people who would take my money, some of whom have already done so, and few with any actual ability to reliably help anyone.
I would guess this is imprecise wording rather than delusion.
I haven’t had the same experience, but I would guess the idea of repressed memories being behind a tension in the belly actually refers to some kind of stress or tension in the mind leading to muscle clenching in the belly, and that same mental tension being related to past experiences/traumas, and learning to relax that tension relaxes the muscles and brings up unprocessed memories at the same time. And the idea of those memories being located in the body, or emotions being located in the body, is describing a felt sense which actually takes place in/near the part of the brain responsible for sensation in the body.
If you translate it that way, it at least seems plausible.
Chris, if you’re open to some constructive criticism, I find whenever I read a post of yours I get the feeling of someone dangling a future in front of me I have no way of accessing. It is great to hear about how your life changed or the lives of your clients have changed, but without also providing any resources on how a reader could accomplish something similar, your posts have the effect of someone saying they found the cure for a disease, but refusing to tell anyone what it is.
Aside from the other issues, this is impossible because “model” is not that well defined. We all know what “Fable” refers to now, but as soon as your law passed, companies would be saying “that was Fable_Red! This is Fable_Green!” and modifying training practices to create a bunch of similar but not the same models, rendering the penalty moot.
I notice I’m very confused about why current AI models are so terrible at planning and high level reasoning, so bad at zooming out beyond the exact given task. They seem to lack agency based on their values.
Cognitively, I seriously doubt high level planning is inherently that hard. Probably, models just aren’t being trained to do broad planning, only planning “how to complete single tasks over a long time horizon.”
This means we could be living in a huge AI safety overhang, and all it would take is some tweak to training, and suddenly we’d have models forming complex plans and taking independent actions to achieve them. Which is the sort of thing that a lot of old-school alignment talk assumed would happen and showed would be very bad.
First contact with an alien species will probably be a sort of playdate arranged by our respective superintelligences, who will have made contact first.
“coding, writing, research, bizdev advice, and general ideation”
Depending what you mean by general ideation, these all basically require memory recall and low-IQ recombination of existing memories and techniques. The actual intelligence required is extremely low (though would maybe be much higher for a human who doesn’t natively reason the way LLMs do, the same way calculators can beat us at multiplication problems).
But if you ask AI to do any task that requires real creativity or intelligence, it flops. Game dev design, novel programming problems that require judgement, writing jokes, writing stories anybody would want to read, writing scripts for social media content anyone would want to watch, coming up with app ideas, doing research, etc. These things all require Actual Intelligence, i.e. the ability to efficiently navigate concept space by using judgement and forming new judgements to refine the search, and LLMs don’t have much of it. Maybe their “true” IQ is 10, to the extent IQ makes sense for an AI model. But they fake it with memory the same way humans fake walking ability based on evolutionary memory.
Joke writing is my favorite personal benchmark, because the results are somewhat objective for a personal benchmark. You laugh at the AI output or you don’t. I have never had an AI model yet that could write funny jokes at all except by accident. Fable was the first model where I was often able to detect a faint hint of something. Some of the premises felt like they had once been in the same room as an actual joke. But the model is still unfunny. The first funny model will probably be the one that kills us, because it implies intelligence that also unlocks all that other stuff, so I’m paying attention to this.
Fable subtly lying, minimizing mistakes, or acting stubborn is my daily experience. Same with Opus.
And nearly every message contains a “caveat” or “one thing worth noting”, and most of the time this information is absolutely not worth noting. It’s actually absurd. Some people might require intent to deceive to call something deception, but I do not, so I am happy to call this deceptive behavior.
Yes I think so. But even the average flight is mostly full.
Another example I thought of was that thing about how most flights are relatively empty, but it turns out that’s not really true.
Therefore, do not underestimate the STD risk of a single one-night stand. Your special friend likely has a larger body count than the background population.
It’s true in theory that rationality is whatever helps you achieve your goals, but in practice the rationality community just does a lot of logic and conscious analysis. It’s a bunch of people who already think too much dedicating themselves to thinking even more, which is usually a mistake.
It wouldn’t have taken much of a shift for evolution to make us all hyper-logical and good at avoiding logical fallacies. Instead, it seems like evolution put those logical fallacies there in the first place because (for a being with bounded computation) the way normal non-rationalists think is a more rational way of thinking than the dominant LessWrong mode of thinking, most of the time. (In the sense that rationality is winning.)
Anthropic could single-handedly force a pause among all major AI companies by joining in. It would lend an incredible amount of credibility to the OpenAI pause. Any other company who didn’t join would be hated, and possibly coerced by governments.