i mean, plausibly i could also make a python program that matches the behavior of a trained model, though this would be quite a bit more difficult. what’s the size multiple at which you’re indifferent between a network-matching program, and a more capable but non-network-matching program?
leogao
what if it turned out that actually the computations needed to reach N are not so abstract as to be incomprehensible? what if we could write out a report explaining all of the specific statistical correlations used in the production of this python program?
tbc, there is no guarantee that the python program matches what any real transformer does in any way.
also, i don’t care that much about advancing linguistics. i mostly care about the ai safety implications
what are some questions you’re curious to answer given such an object?
i find it hard to overemphasize how much my research workflow now is different from a year ago. it feels like almost stepping into an entirely different job. i feel like someone who has spent a decade punching COBOL punch cards and then suddenly been given a lisp machine.
don’t have a super good formalized definition. mostly know it when you see it. but clearly vastly more interpretable than a normal model of the same level capability
suppose i could create a human-understandable python program that is equivalent in performance to an N parameter language model. how big would you expect N to have to be for it to be a very interesting object to study? 1M? 10M 100M? 1B?
it also seems possible to discover novel attacks that transfer across multiple models. eg original GCG was this
what’s the strongest case for why something IDA-like cannot be aligned? my current best story for how ambitious mechinterp might be possible depends on an IDA-like assumption. it would be good to know now if this is a doomed endeavor.
what are the strongest reasons that solving interp might be net bad for the world? some ideas to get started, but i’d like to get all of your takes.
maybe it actually advances capabilities more than it helps alignment
maybe it isn’t actually as useful as we think, but people use it as an excuse to scale to dangerous levels
oh sure. if you’ve already decided you want to quit, then you should just quit.
the hypothesis is that people who refuse to work on capabilities will be given no power, even if they are otherwise treated well. but if you are at an ai lab and have no power, should you conclude that this is because you fundamentally can never gain power (your hypothesis), or that in general gaining power requires you to be pretty competent in ways that are rare and therefore most people will have no power anyways (the null hypothesis)
your time is probably worth more than your monetary donation.
how would you distinguish this from a null hypothesis of gaining influence inside companies being hard in general for mundane reasons?
my timelines have in general gotten shorter over the past year. the biggest update for me was actually codex being absurdly good at running experiments. i don’t know enough math to know how impressive the math things are, except by deferring to other people.
no, obviously openai doesn’t prohibit employees from talking to therapists
my understanding is the proof is very different anyhow. also, the model is extremely good compared to astra at resolving many other open math problems too, which should give some evidence for this model simply being extraordinarily good.
i don’t know about the details of sebastien and buckmaster and levent’s conversation. but i don’t think openai as an organization has done anything obviously untoward here?
i mean, there’s always the boring python program (just write out all the weights). but assume that it’s not like that, and looking at the code straightforwardly tells you basically anything you want to know