sure the evidence is less strong than if every model had linear increases in its persuasive ability, but like, the stock market still goes up? if we zoom out at all the progress is very fast and very clear.
Beckeck
models also out persuade national championships debaters and professional canvassers, although the models regress when their ‘quantity of facts presented’ is limited to the human norms.
It’s covered in the first dispatch here.
https://aistop.watch/p/shifting-perspectives
They aren’t airgapped because air gapping is expensive (mostly in lost productivity) and annoying, and insufficient in and of itself (and probably not the most pressing security failing),
and the companies / the employees don’t wanna, they wanna race.
I appreciate your work on many issues, but find myself pretty strongly disagreeing.
finding minor criticism of a sub point that is technically valid (often only if you interpret things noncentrally) doesn’t mean the broader claim is wrong.
Said imposed more cost then provided value, and refused to cooperate with extensive efforts to find third options. Banning that seems productive.
What are ways that y’all avoid/manage goodharting?
Do you have a process to change goals if getting politician’s signatures seems less valuable?
The initial release is just in the US. It might get an international release if sufficiently successful, but we don’t know. Similarly we don’t know about streaming, but i’d bet it will get onto some platform eventually.
“The AI Doc” is coming out March 26
Also have this issue on galax s24. (And not on other parts of the website)
Hope this goes well!
random pitch but—maybe add anki integrations for extra nutritious content?
Yes, see PauseAI (even if I disagree with some of their positions, i’m glad they exist and hope that soon there exist multiple such orgs (but don’t donate to StopAI, they don’t appear serious imo))
upvoted for topic importance.
thanks, I appreciate the reply.
It sounds like I have somewhat wider error bars but mostly agree on everything but the last sentence, where I think it’s plausibly but not certainly less worrying.
If you felt like you had crisp reasons why you’re less worried, I’d be happy to hear them, but only if it feels positive for you to produce such a thing.
we might disagree some. I think the original comment is pointing at the (reasonable as far i can tell) claim that oracular AI can have agent like qualities if it produces plans that people follow
yeah, if the system is trying to do things I agree it’s (at least a proto) agent. My point is that creation happens in lots of places with respect to an LLM, and it’s not implausible that use steps (hell even sufficiently advanced prompt engineering) can effect agency in a system, particularly as capabilities continue to advance.
“Seems mistaken to think that the way you use a model is what determines whether or not it’s an agent. It’s surely determined by how you train it?”
---> Nah, pre training, fine tuning, scaffolding and especially RL seem like they all affect it. Currently scaffolding only gets you shitty agents, but it at least sorta works
Top post claims that while principle one (seek broad accountability) mightbe useful in a more perfect world, but that here in reality it doesn’t work great.
Reasons include that the pressure to be held in high standards by the Public tend to cause orgs to Do PR, rather then speak truth.
know ” sentence needs an ending
“ARC (they just changed names to METR, but I will call them ARC for this post)”—almost but not quite—
ARC Evals (the evaluation of frontier models people, led by Beth Barnes, with Paul on board/ advising) has become METR, ARC (alignment research center, doing big brain math and heuristic arguments, led by Paul) remains ARC.
″ football, hockey, rugby, boxing, kick-boxing and MMA to be amongst the worst sports for this stuff.” - - I’m not up to date on the current literature but I’m pretty sure this list is rather wrong. I don’t remember all the details of the study I do remember (and I don’t have time for a lit review) but in it women’s high school soccer actually had the highest concussion rate (idk if it was per participant season or hour or per game minute or...).
I reject the paraphrased claim in at least a couple ways:
One, deontological or virtue ethics induce people to make bad choices when it comes to trolley type problems, compared to consequentialists who make the better choice(s).
Two, it begs the question by presupposing evil people instead of fallible humans. I expect this is downstream of broader mistake v conflict theory stuff, and don’t expect us to resolve ethics in this comment section, but seems worth calling out foundational disagreements as being sufficient to explain differing conclusions.