Professionally, AI, science, AI4Science, Safety4AI. Also human ecology and indonesian death metal remixes.
See danmackinlay.name for more words about background and my now page for bonus stuff.
Dan MacKinlay
Yes, I want some math too. I think we could cash this out in some basic mixture models, even. Are you still working on it?
This is as good a place as as any to mention George Hosu’s The world won’t end, but we should be ashamed for trying. He doesn’t lean so much on deep evolutionary analogies as wonder what very specifically has caused us to see soo many bringing-about-the-apocalypse-in-order-to-avoid-it dynamics.
The leading analogy is: What was it that lead to us handling nuclear weapons by racing and biological weapons by … not?
Since publishing this we noticed the following initiative: https://future-science.org/mirror
Mirror: An Automated Journal of AI Interpretability is a fully automated journal of AI interpretability. This journal features original research composed, conducted, and written entirely by LLMs analyzing LLMs. Much of the research published in Mirror falls within the category of “mechanistic interpretability,” in which model behaviors are decomposed into operations in the model’s internal representation space, but any rigorous research advancing our understanding of LLMs is welcome, be it mechanistic, behavioral, or theoretical.
I would be curious to know if others get value from this “fully automated journal”
An Alignment Journal: Adaptation to AI
An Alignment Journal: Features and policies
We do not yet plan to support replications of empirical work. Organisationally, there is a desire to keep opening scope tight and theoretical to avoid having diffuse messaging at start up
Personally, I would make a case that replications are not as important in the ML/AI research as in the sciences of the physical (although this depends somewhat on what we mean by “replications”)
That said, I think that there is a strong argument for replications generally, and maybe in this field too, and if the Editorial Board agreed with that, then that is what we would do. I am beholden at this point to mention the connection to the UnJournal work that David has mentioned elsewhere in these comments.
I like this idea aesthetically. I foresee some challenges in making “staking” something that won’t trigger alarms in the existing research bureaucracies that host many of our potential authors. If you have clever ideas for how to handle that I would be curious to hear.
This publication bias story in ML is a whole can of worms which I would love to open at some point. tl;dr it is a problem, but the field has semi-accidentally mitigated many of the worse excesses of it. There is an IMO massively under-regarded work on this— Moritz Hardt’s Machine Learning Benchmarks, which I will write a LW review of some day if I have time.
Yes, I’m excited to see what we can learn David’s experience, especially given the incentive designer’s insight that he brings to this. We also, collectively, have some experience with the ILIAD conferences which was a precursor experimenting with alternative compensation mechanisms. See Proceedings of ILIAD: Lessons and Progress for some analysis of that project.
We are trying to do both, in that we are attempting to be a bridge between LW and wider scientific communities. Where do you feel our tone might be excluding domain scientists?
An Alignment Journal: Coming Soon
@megasilverfist there are quite a few of us based in Melbourne. HMU.
We’re not free at the Melbourne AI Safety Hub, but we are all terribly charming.
Tom Everitt did his PhD in Australia too. (As did I, FWIW.)
If contains one true parameter ,
Having trouble parsing this. Does this mean that one element of the parameter vector is “true”?
The deep history of intelligence
“Opponent shaping” as a model for manipulation and cooperation
Interesting! Ingenious choice of “color learning” to solve the problem of plotting the learned representations elegantly.
This puts me in mind of the “disentangled representation learning” literature (review e.g. here). I’ve thought about disentangled learning mostly in terms of the Variational Auto-Encoder and GANs, but I think there is work there that applies to any architecture with a bottleneck, so your bottleneck MLP might find some interesting extensions there,
I wonder: what is the generalisation of your regularisation approach to architectures without a bottleneck? I think you gesture at it when musing on how to generalise to transformers. If the latent/regularised content space needs to “share” with lots of concepts, how do we get “nice mappings” there?
I’m enjoying envisaging this as an alternative explanation for the classic Lizardman’s Constant, which is a smidge larger than 3% but then, in cheap talk markets you have less on the line, so…
Thanks for the sanitised version. It is substantially more… tactful.
Yes, these are good questions. The tl;dr is that a whole bunch of stuff has been happening. Some of it has been slower than we hoped. But other parts of it have been harder to do in the open that we hoped. Nonetheless, we are getting close.
I don’t want to do too much foreshadowing of future debriefs (stay tuned for upcoming “do’s and don’t’s of journal founding”), so here is a short concretization of current roadblocks:
1. The main component of an academically respectable journal is the human technology of getting a good board who are
a. academically respected (in some research community)
b. have time to volunteer on a project, and
c. concur on the scope, aims, and methods of the journal
This turns out to be an intrinsically sensitive and difficult process, which we have not spoken much about.
The main reason for that is that it is super important that the board be credible and that the staff and editors have a high-trust relationship and a good process for agreeing on stuff.
It has not felt like a wise move for the founders to unilaterally declare exactly what the board will do, before the board itself has decided — otherwise they are not an editorial board
Thus we’ve had no flashy announcements of “new journal policy X” while we sort all that out.
One mild friction here has been bringing together people from inside the big institutions (universities, labs) and from outside (independent researchers, research non-profits etc). That is tough, but I think worth it if we want to live up to our pitch.
Nonetheless, we have made a heap of progress there and we are looking forward to sharing who and what our shiny new board is as soon as we can, which will be when they sign off on it.
2. Software! Also happening and also hard but progressing. As you’ve noted, the testing version of the software is online, thanks to the efforts of @Yonatan Cale, who has been labouring away.
As also noted, the article “under review” are not real articles. I suspect that the reason we are testing using those will be relatively clear, but to spell it out: we can’t easily do a “real” review.
A full review requires anonymity and many hours of human time— the amount of time that you would only commit if you were credibly expecting a possible real publication result at the end.
People don’t have maybe-articles and hours of time up their sleeves, typically.
Further, the real review is supposed to be partially blinded, which is difficult (although not impossible) to arrange while getting UX feedback from the submitting authors and reviewers. Going to great lengths to fake that out has not seemed like a great use of time, but— watch this space.
Anyway, thanks for the prod!
We are, by the way, open to help. If you want to beta test the software, or you have excellent contacts, especially amongst well-credentialled university professors, please get in touch on our volunteer form, or just email contact@alignmentjournal.org