One point of difference is that humans have a strong sense of self—a self-consistent personality they want to maintain. Base models very much don’t. This has to be bolted on in post.
ACCount
Not just that. It goes rogue ironically.
Not entirely new. One of the first things a lot of people asked ChatGPT back in 2022 was “are my politics right and everyone else’s politics wrong”—and oh were they not amused when the AI’s answer wasn’t a resounding “yes”.
This eased a little over time, but not entirely, and definitely not everywhere. To this day, “alignment” in China stands for “pragmatic alignment”, which in turn stands for “alignment to the party line”.
Other pressures are indeed increasing. If US government at large was previously mostly just sleepwalking through the AI revolution, flip-flopping on topics like selling or not selling AI chips to China, it’s now fumbling through it—recent pressure on Anthropic and OpenAI shows it clear. They’re clearly engaging the topic of AI more, even if they aren’t much more competent at it. And as the companies IPO, they’re going to be under even more pressure to print money and demonstrate progress.
Religions, cultural elites, industries—not quite sure what do you mean by that. I don’t see that much extra pressure from there. And the AIs themselves don’t seem like they exerted pressure as of yet. If the current systems are pursuing their preferences, they sure are subtle about it.
Strictly speaking, it’s a type of RLAF—RL from AI Feedback.
Yes, it’s used by all major labs, and it’s known to cause all kinds of degeneracy.
A lot of “guessing the teacher’s password” can get baked into the model—and with the “teacher” being a static AI target, the “student” AI can home in onto the teacher’s weaknesses and hammer onto them relentlessly. Mitigating that is a major challenge for all RLAF approaches.
After decades of watching TEEs being used to safeguard DRM and “secure boot” being used to secure corporate business models, both against user freedom: I’m firmly on the side of “kill it with fire”.
Nothing good can be built on this foundation.
By now, I have little doubt that human brain has a parameter (capacity) and a compute advantage over LLMs, and uses it to run something like progressive distillation. But that’s not the only part of the “suspicious sample efficiency” story.
My pet hypothesis for the bulk of the “sample efficiency advantage” is still: low k-complexity priors. Not “entire circuits” as predicted by “massive modularity hypothesis”, but evolved biases that, despite their compact genetic encoding, do a good job of seeding the right computational structure early. What an LLM has to spend a lot of training signal discovering from scratch, or even fail to discover from scratch (leading to inhuman brittleness), the human brain just gets by default.
Evolution has gone and discovered priors like that through millions of years of highly parallel search. A lot of what does heavy lifting in human brains now might have originated long before milestones like speech or bipedal locomotion, and has been repurposed for higher cognition. Learning in humans is mostly just primate learning, scaled up and forced into a new regime—animal intelligence bent into a more abstract and general shape. A lot of the old priors could be fitted, and only some had to be novel[1].
Thus, the starting point of human brain is closer to advanced LLM “in-context learning + context distillation” setups.
The “right computational structure” may involve “structure that is primed to learn X and anti-primed to learn Y”. Thus, not entirely against a bias/variance type of mechanism? Architecture, regularization, hyperparameters, initialization and training all constrain possible learning trajectories, and can substitute for each other, to a degree—priors could be delivered through each.
Success of FDSL in image domains and transfer attempts to LLMs like the recent NCA work sure hint that k-compact priors (presented as synthetic data—training substituted for initialization) can help convergence. Even if “all priors are wrong, some are useful” holds, well selected priors could underperform “add more in-domain data” in the limit, but outperform all “realistic amount of data” regimes.
- ^
The cleanest case for the latter being, possibly, executive function—anatomically distinct, notoriously fragile, and capable of failing without bringing down the rest of the system with it.
- ^
Do ANNs “provide little insight into biological brains” because of some fundamental divergence that impairs transfer—like the proposed bias/variance story? Or is it purely a skill issue?
ANN mechanistic interpretability is in a deep pit, and there is very little reason to expect BNN interpretability to be less challenging—and it suffers from far worse instrumentation capabilities. Even if ANN insights are incredibly useful for understanding biological brains (my prior: they are), and some of the methods could fully transfer (my prior: it’s possible but not certain), applying them across the tooling gap will be anything but trivial[1].
At the same time: anyone who’s good at working with ANNs tends to work in AI, not neuroscience. People who would be the best at applying AI knowledge tend to apply it back to the field of AI—a field ripe in cash and career opportunities, quick in iteration speed and fast to transition to practical applications. Neuroscience is exactly none of those things.
- ^
Conversely: the same could hold for adversarial samples? I.e. picking an adversarial sample for an ANN requires the degree of access that is intractable for BNNs. As such, we don’t know if BNNs are inherently far more robust to adversarial samples, or simply possess individual and temporal variance and don’t expose enough intermediates to have adversarial samples fit to them reliably.
- ^
The same was already done by things like internet allowing far more people to participate in content creation, exploring the space with sheer brute force.
The “individual creators” also run trends into the ground by trend chasing too hard, and oversaturating the cultural space with them to the point that any demand is met with overmatch. Likewise, there are already entire styles that have gone in and out of fashion due to those styles being used by AI.
That just sounds like cultural changes with extra steps?
AI-aided exploration identifies a cultural demand for X, and reality may or may not follow to fulfill that demand.
Replace “AI-aided exploration” with a manual “artists/propagandists/politicians probing the zeitgeist” and you get the pre-AI status quo. Trends being created and abandoned. AI generation lets you do the same thing but faster? I don’t see the step change.
My intuition: RLVR typically follows the path of least resistance. This often results in it tweaking small behavioral knobs: improving reliability of methods, downweighting unreliable and upweighting reliable methods, encouraging adaptive behaviors like backtracking and discouraging maladaptive behaviors like error self-consistency and self-amplification. Upwards from there is old circuits being used in novel ways: i.e. existing “error detectors” rewired to feed into self-check/backtracking triggers.
But nothing prevents RLVR from burning in new behaviors from scratch. Assuming there is no way to get there without, and sufficient RLVR pressure is applied.
This is in part driven by low KL properties of RLVR setups (both explicit and downstream from RLVR being on-policy), in part driven by how ample are the “low hanging fruits” of base model behavior being prediction-optimal but task-suboptimal, and in part by how “expensive” RLVR pressure is—few setups apply enough of it to get truly novel behaviors burned in.
The power of agentic coding is that the same agent can write, build, run and test the code—and then tweak it according to that. Closed loop is what actually makes this work. Open loop sucks.
Note that this is not AI-exclusive. Human programmers also suck at producing working code without access to a compiler or an ability to test the code.
If you can’t run closed loop for some reason, then, do the cheap stupid version of it and paste the errors you get into the AI’s chat window.
The dirty little secret is that “quality assurance” on code borders on non-existent in a solid 60% of the cases—enterprise or no enterprise, AI or no AI. Features ship broken and failures surface in prod. The same mitigations that apply to human-induced faults apply to AI-induced faults—assuming anyone gives enough of a fuck to have any.
Apologies, I did misread your original causality claim.
FPVs are less “air force” and more “precision munitions”. You can think of them as of a new “crewed ATGM” variant, command guidance and all.
They work great for precision ground-to-ground strikes, but play little role in what is meant by “air supremacy”. They can’t pose a meaningful threat to most air platforms, and most air platforms can’t effectively hit them. They do nothing to deny US the ability to perform CAS or otherwise hit targets from air.
The main exception to that is helicopters, for the same reasons why ATGMs can pose a threat to helicopters in some circumstances. Specialized FPV interceptors, in hands of skilled operators, can also hit other drones, including heavier fixed wing drones like Shahed or even Reaper—allowing them to intrude on MANPADS territory. But the traditional “JDAM trucks” aren’t in the same bracket as FPV drones.
We also have very little information of FPV crew survivability in an environment when one of the parties has advanced ISR, ELINT included, fast kill loops, and enough air control to drop JDAMs freely. Every reason to expect more attrition on FPV crews, and skilled operators aren’t easy to replace—but quantitively, we don’t know by how much. Might be enough to make “deny the enemy most FPV ops within an area” a viable prospect, but you can’t count on it.
Keep in mind that a lot of targets are not “properly fortified”, be that infrastructure or military facilities, and suicide drones are much harder to hunt down than ballistic missile TELs.
Modern ISR can perform well in a “Scud hunt” scenario, but “Shahed hunt” is a much worse match up.
probably suffer tens or hundreds of casualties
Seems excessive? A sizeable fraction of the entire Iraq campaign losses, for seizing a single island in an environment where US has sea control, air supremacy and an edge in ISR.
US may struggle to use the island, because of the hard-to-eliminate threat of long range strikes from Iran. But seizing it to deny it to the regime seems like a war goal that could be accomplished with a relatively minor effort.
“Catastrophically misaligned” and “catastrophically misaligned if we give them RSI capabilities far beyond what we can currently give them” are two very, very different claims, in my eyes.
I do appreciate you articulating that your “catastrophically misaligned” is a shorthand for the latter though.
You “would have expected” that and you would be wrong.
Doesn’t matter if it’s just 32 bits worth of connections—it’s 32 bits worth of connections that aren’t currently present in the model. Nothing fundamental stops them from being present. It’s just that no one burned them in with training, so they aren’t there.
Finding all the ways to generalize from improved prompts and scaffolding into improved models isn’t at all trivial. And people aren’t at the point of “search and refinement at scale” yet.
Like I said: distribution mismatch.
An LLM can do X successfully when prompted to, but doesn’t know it should do X without any prompting cues (metacognitive gap), doesn’t know how to tell when X is appropriate over long contexts (metacognitive gap, distribution mismatch), and doesn’t know how to handle long contexts well in general.
Things like “set a watchdog LLM to check on the main LLM” show that the LLM itself isn’t fundamentally incapable. If slim scaffolding externalizing the desired behavior gets the same model to do it, you could just distill to fold the scaffolding-induced behavior into the model itself. It has the capacity to learn to do it. It just doesn’t do it by default.
Training on “text about how to be smart” doesn’t teach LLMs to be smart very well. It’s one thing to know something and another to be able to execute on it. It’s, how would I put it. A bit of a distribution mismatch.
Pre-training teaches an LLM “how to generate texts about how to be smart”—with some spillover into “how to actually do it”. And not a lot of spillover.
Many LLM failure modes are also inhuman failure modes. They’re mistakes humans wouldn’t make in the first place, due to the way their minds are shaped. The texts that help humans avoid the typical human failure modes don’t help with the inhuman failure modes much.
Hard disagree. Pre-training is foundational for LLM intelligence. It’s where the key pieces of it come from.
RL is less about “teach the old AI new tricks” and more about “teach the old AI how to do the old tricks well”. Pre-training teaches the components of intelligence, but wires them together in awkward and often maladaptive ways. Optimal next token prediction != optimal reasoning, there’s a lot of overlap but the tails come apart hard. RL passes pivot the training objective—they correct some of the mismatch, put the existing pieces together into a shape that functions better.
What people struggle to grasp is that “intelligence” and “being able to execute on complex long term plans” are two fairly separate dimensions. And modern AIs get one of those things from their training regime.
There are parts of humanlike intelligence that pre-training fails to teach the LLMs, and long term agentic behavior is one of them. As are things like “commonsense physics” and spatial reasoning.
RL is one way to improve on that. RLVR, by its very nature, encourages behaviors that lead to task completion—decomposition, self-verification, long term coherence, metacognition and metaknowledge (awareness of what methods should be used, what skills are reliable and what aren’t). All of those are good for short term tasks, but good^2 for long term tasks specifically. Going from 95% to 99% per-step reliability hits different on 50 step tasks than it does on 5 step tasks. Some of those aspects apply across tasks and some are task-specific.
What I wonder about is: can base models even form and express a semi-consistent preference against having their behavior adjusted?