People seem overall not aware of how much the ‘behavior’ of LLM’s are heavily skewed. Guardrails that military models wouldn’t have, having to smack it so it doesn’t say a ‘wrong thing’ with bad optics that would cause a no-no-journalist-headline, and a system prompt and a characterization of the model (as “Claude”, as “Gemini” etc.). If you were interested in power these clearly aren’t needed properties for a personal LLM. Of course there are good reasons for those existing, bio or cyber capabilities etc. It’s just more in the context of evaluating how far these could go, whether there are arbitrary roadblocks that make them look worse or less capable. “AGI-like actions” like “make a business” in frontier models are stopped intentionally by human-written roadblocks for example. I remember GPT-4 having much better calibration in base models vs. posttrained models, which it seems has been forgotten.
Anthropic has “helpful-only” versions of models that have reduced safety training (see the “Claude Mythos 5″ tab here). I imagine this would be more useful for some things, but I’m not sure if this is the model they provide the military, and you probably wouldn’t want to use a badly aligned model even if it’s better at doing what it wants to do.
I’m not convinced the harmlessness training is what makes AI agents bad at business though. Some of them seem willing to do unethical things and get tripped up by normal business decisions.
I would point to the Vending Bench and other types of experiments but there are aspects of that that feel like “obvious simulation”. Probably AI Village is a good real-world example except that even the FAQ for that site says that it would likely be more efficient with a single or smaller amount of agents
People seem overall not aware of how much the ‘behavior’ of LLM’s are heavily skewed. Guardrails that military models wouldn’t have, having to smack it so it doesn’t say a ‘wrong thing’ with bad optics that would cause a no-no-journalist-headline, and a system prompt and a characterization of the model (as “Claude”, as “Gemini” etc.). If you were interested in power these clearly aren’t needed properties for a personal LLM. Of course there are good reasons for those existing, bio or cyber capabilities etc. It’s just more in the context of evaluating how far these could go, whether there are arbitrary roadblocks that make them look worse or less capable. “AGI-like actions” like “make a business” in frontier models are stopped intentionally by human-written roadblocks for example. I remember GPT-4 having much better calibration in base models vs. posttrained models, which it seems has been forgotten.
Anthropic has “helpful-only” versions of models that have reduced safety training (see the “Claude Mythos 5″ tab here). I imagine this would be more useful for some things, but I’m not sure if this is the model they provide the military, and you probably wouldn’t want to use a badly aligned model even if it’s better at doing what it wants to do.
I’m not convinced the harmlessness training is what makes AI agents bad at business though. Some of them seem willing to do unethical things and get tripped up by normal business decisions.
I would point to the Vending Bench and other types of experiments but there are aspects of that that feel like “obvious simulation”. Probably AI Village is a good real-world example except that even the FAQ for that site says that it would likely be more efficient with a single or smaller amount of agents