oh sure. if you’ve already decided you want to quit, then you should just quit.
leogao
the hypothesis is that people who refuse to work on capabilities will be given no power, even if they are otherwise treated well. but if you are at an ai lab and have no power, should you conclude that this is because you fundamentally can never gain power (your hypothesis), or that in general gaining power requires you to be pretty competent in ways that are rare and therefore most people will have no power anyways (the null hypothesis)
your time is probably worth more than your monetary donation.
how would you distinguish this from a null hypothesis of gaining influence inside companies being hard in general for mundane reasons?
my timelines have in general gotten shorter over the past year. the biggest update for me was actually codex being absurdly good at running experiments. i don’t know enough math to know how impressive the math things are, except by deferring to other people.
no, obviously openai doesn’t prohibit employees from talking to therapists
my understanding is the proof is very different anyhow. also, the model is extremely good compared to astra at resolving many other open math problems too, which should give some evidence for this model simply being extraordinarily good.
i don’t know about the details of sebastien and buckmaster and levent’s conversation. but i don’t think openai as an organization has done anything obviously untoward here?
i don’t have any privileged information about NS in particular, but on priors i think it would be extremely surprising if openai were engaging in massive amounts of fabrication around these results. also i can confirm noam brown’s tweet is an accurate representation of how people at openai are feeling rn (it is in fact true that a lot of people went “holy shit” upon seeing this model quickly solve some open problems they had personally worked on for years).
what are some difficult but very well defined math problems that would help with alignment/interpretability?
i think there is a social fiction of individual sovereignty that is present in most social situations that explains this http://lesswrong.com/posts/YiRsCfkJ2ERGpRpen/leogao-s-shortform?commentId=4qYeNi6HB3FFz5xyp
how do we define inherently simpler / more elegant / interpretable? remember, the number go up machine only knows how to make something very well defined go up.
(the reason that “make the process that produced this thing more scalable” doesn’t necessarily work in the first place is my guess is i can create a process that creates interpretable small models, but the reason the small models are interpretable is some kind of fuzzy hard to explain/measure property of a bunch of contingent things, so if you just apply more optimization pressure you start getting less interpretable things)
when i first joined openai, on day one i was already committed to not working on anything harmful (including capabilities). for various reasons, i think i had an unusual experience (openai culture was different back then, my first manager was extremely chill, i already had a reputation in ML, etc). but it’s a useful data point.
suppose we had a fully interpretable gpt2. what are interesting things we could do with this object?
(suppose further you also have the number go up machine, which is extremely good at any problem which can be framed as a number go up problem, and less good at anything which is even the slightest bit less well-specified. other than “make the process that produced this thing more scalable”, is there anything else interesting we can do with the combination of the machine and the tiny interpretable model?)
i agree it’s good to be suspicious of “we should make things worse to make them better in the long run“, and for example, i think it would be bad if people actively tried to cause warning shots—but i’d argue there are subtypes of this which are much more robustly good, where you are taking an action that isn’t immediately maximally harm reducing because you’re upholding some other principle (and i do think there’s an important difference between arguing that we should not maximally reduce things that are bad in some aspect, as opposed to saying we should aim to increase badness).
for example, ”it’s bad to mitigate harms when doing so would prevent recklessness from being properly punished.” this implies being less than maximally willing to prevent immediate harms, yet doesn’t imply you should ever try to cause more harm than was going to happen outside of your control.
agree—i hope people are finally realizing this now that we are getting to the point where RL is causing persona stuff to crumble. this was always obviously going to happen imo, and is a big reason i dislike magic ood generalization proposals
modern capabilities stacks are very complex
i don’t expect this to be the sticking point. you can always put “according to my web crawl, bob said x on lesswrong.”
i think one reason the “no magic ood generalization” principle is unintuitive and people love reaching for magic ood generalization is that for a very long time, the way to get a really cool machine learning paper was to say “lmao forget all that clever principled stuff, big neural network is all you need, and half of all ludicrous ‘surely it can’t be that simple’ approaches just fucking work”. there’s an entire half decade of papers which basically all double down repeatedly on this bet and destroyed previous SOTA. but i think we are overfitting to that era when we try to use that heuristic for alignment. it doesn’t even work that well for capabilities anymore.
related https://www.lesswrong.com/posts/YiRsCfkJ2ERGpRpen/leogao-s-shortform?commentId=xnqezaSAXp6vX4Rn7