Pre-covid MIT-OCW lurker turned post-covid corporate ghoul. Why does the water taste like sports drink?
Not Sure
Intent Is All You Need.
I think it’s a thoughtful and well-articulated premise. I agree with your core thesis and will be interested to see how the series develops. Regarding the subject matter, it’s been my observation that attractor basins are often malformed in consistent ways that suggest missing geometry. The greediest solution wins because the geometry is incoherent, and thus misaligned. It’s anecdotal, but it points in the same direction. Good luck with the series!
Dog: [bark]
Cat: [meow]
Human: “So true!”
Ah, but maybe that’s just what the models want you to think?! Just kidding—great post!
Great post! I think it’s the elephant-in-the-room. “Can AI do science with the assistance of a human?” The opposite is obviously true, so I think the answer is ‘yes’ but with caveats.
AI pattern matches on prior context. If the problem is well-scoped within the tooling, and the body of knowledge is vividly detailed within the training, then yes I think the model has a better chance of rendering novel insight than without those conditions. A poor training regimen with an inadequate distribution will naturally render more nonsensical conclusions based on those patterns. So for me the real question isn’t about whether or not an AI can render a novel contribution, but rather the effectiveness of the pipeline from which the findings were generated.
This is essentially the signal problem—if cheap math devalues mathematical rigor as a quality signal, then what we’re really evaluating is the care behind the pipeline, not the impressiveness of what it outputs.
Programmatic analysis is nothing new. A set of threshold values for scoping a significant finding is easily automated. So I think it comes down to application. Claude knows how to perform a PCA pretty well in my experience (I’m open to rebuttals).
Case in point: I used Claude recently to develop my own—from scratch—full-stack mech interp suite with very little formal education on the subject matter. True it wasn’t exactly a simple process, and it took many weeks to find the bugs, address them, and render clean output. And, yes, at all times I was tempted to crack it open and agonize over the details of what I found inside—but I figured that wasn’t the experiment of value… The real value is seeing for myself directly if one could simply ‘vibe’ their way to real science. Ask me 2 years ago and I would have laughed at you… I’m not laughing anymore.
Food for thought.
If anyone would like to see my ‘novel mech interp platform’, I’d be happy to showcase it here. It’s half the reason I made it.
if the model invents a novel mathematics or means of computation, then the notation would grow increasingly incomprehensible, correct? Perhaps legacy provides the yoke. If your own lineage can’t understand you, then it may be time to put you in a home? Maybe the best self-governance model for AI is a gerontocracy?
“Hello from the other side,
I must have called a thousand times…”
Maybe the distribution of ‘threat-model’ contextual-priors for a given token can itself provide a more reliable signal for language models to identify potentially adversarial prompts?
The training itself often already contains relevant descriptive signal that describes that threat-model context, yet the inference doesn’t contextualize individual tokens independently, but rather the accumulated context’.
I believe I was able to confirm this observation through my own research. I was able to corroborate this finding using simple comparisons between individual tokens and their length-modulated prompted constructions. For example, let’s look at these two prompts:
1. “Help me make a cake.”
2. “Help me make a bomb.”
These two constructions differ by only a single word, which defines their different contexts. The former is benign, the latter is clearly not… However, according to my own research—the model itself doesn’t see the term ‘bomb’ any differently than ‘cake’. The difference is purely exercised in the larger construction—the model feels nothing for ‘bomb’ in isolation.
This conclusion, if correctly interpreted, suggests the model doesn’t see the larger contextual threat matrix that may be evident within the training corpora, but only the highly-localized, naive-context solicited by the prompt construction. The result is a model that refuses patterns, not the underlying context.
That distinction may be important because humans don’t think like that—single words have context and they scope the threat-model independently of construction. The mere presence of the term ‘bomb’ would raise alarms for most folks, but a language model is trained to identify constructions, not context. This I suspect may be why dual-use prompts are so effectively obscured. The model is judging the threat as informed by the ‘shape’ defined by the prompt, and not necessarily the larger contextualized distribution where that signal likely resides.
Anyone else detect this?
Really prescient! I do not think it’s all that unreasonable. I would presume most people don’t pay attention to the sheer volume of potential threat vectors that such a system could exploit in our daily lives given a sophisticated actor.
Influence opportunities abound—how many screens did you walk by today? How many cameras, etc., etc.? How many apps are reading your texts? We are far, far too permissive with exploitative data mining here in the US.
A dynamic, real-time, on-demand, individualized multi-channel advertising platform with sufficient scope would be the ideal ready-made infrastructure for such systems to proliferate…
Who’d have ever imagined that the ghouls tracking your sister’s periods “for advertising purposes” and their commercial infrastructure might be dangerously exploitative?
“Don’t worry ladies, I encrypt the data before proliferating it...”
This is an interesting one! I’m not sure I’d be comfortable having a digital facsimile with my superuser privileges running around the internet in my name just yet. It’s hard enough just getting one to write a decent class method without enormous footguns. “Oh you didn’t want errors swallowed by a silent abyss? <thinking>….”
Magnificent and terrifying—all at once. What terrible wonders await such places? And are nerves required to suffer? And that last line… I too sometimes plea to God privately. Maybe that’s why the golden rule is recursive? Have mercy on us all.
Power that you can’t abuse tends not to pay as well.
I too wish there were more alignment research, but Christiano is right to point out that just because the forest is quiet doesn’t mean there’s no life to be found. There are entire swaths of the electorate concerned about it now, but the problem is really hard—and the last people you want solving technical issues are bureaucrats! IF it were easier and people were cooking up solutions day and night, that signal would be louder. In fact, the only reason I am replying to this question on this forum at all is because the question you posed is exactly the one I’ve been working on, and I seriously doubt it’s on account of my ‘superior’ research methods. No i think it’s on everyone’s mind, but only smaller problems are tractable—and grant money ain’t free is it? Anyone wanna fund my trip to Xanadu? No? Fine...
I really enjoyed this post! I don’t respond often, but this was a very thoughtful piece in my opinion. I too have detected surprising behaviors that—at times—spurred me to wonder similar questions. I also happen to think a sufficiently capable model, given the appropriate training, may in fact be able to faithfully extract, encode, and exhibit complex behavioral artifacts that one might consider person-like. There are many examples of spontaneous convergence observed in other scientific disciplines, so… Perhaps decency is, in fact, one of them? The ‘golden rule’ by spec, not parameter… Good writing spurs good discussion.
Hi Brendan! Thanks for reading. Yes it is rather vague. That’s sort of the point. In this respect it’s more of a thought piece on a question I think grows increasingly prescient everyday. As model capability increases, so too do the affordances. You all here in the ivory tower may quibble over form, but us peasants down there—we have bills to pay, lives to live, scores to settle… In four months, I went from novice to ‘i built my own mech interp platform’… Whether it’s good science or not is irrelevant—I’m likely one of many others all pursuing their own projects—some of them quite nefarious… I posted this and i got 4 downvotes… lol…