Cameron Berg

Karma: 730

SERI MATS ’21, Cognitive science @ Yale ‘22, Meta AI Resident ’23, LTFF grantee. Currently doing alignment research @ AE Studio. Very interested in work at the intersection of AI x cognitive science x alignment x philosophy.

There Should Be More Alignment-Driven Startups

Vaniver, Judd Rosenblatt, Cameron Berg and phgubbins

31 May 2024 2:05 UTC

51 points

13 comments11 min readLW link

Cameron Berg 6 May 2024 19:06 UTC
5 points
1
in reply to: Ryan Kidd’s comment on: Key takeaways from our EA and alignment research surveys
Thanks for all these additional datapoints! I’ll try to respond all of your questions in turn:
Did you find that your AIS survey respondents with more AIS experience were significantly more male than newer entrants to the field?
Overall, there don’t appear to be major differences when filtering for amount of alignment experience. When filtering for greater than vs. less than 6 months of experience, it does appear that the ratio looks more like ~5 M:F; at greater than vs. less than 1 year of experience, it looks like ~8 M:F; the others still look like ~9 M:F. Perhaps the changes you see over the past two years at MATS are too recent to be reflected fully in this data, but it does seem like a generally positive signal that you see this ratio changing (given what we discuss in the post).
Has AE Studio considered sponsoring significant bounties or impact markets for scoping promising new AIS research directions?
We definitely want to do everything we can to support increased exploration of neglected approaches—if you have specific ideas here, we’d love to hear them and discuss more! Maybe we can follow up offline on this.
Did survey respondents mention how they proposed making AIS more multidisciplinary? Which established research fields are more needed in the AIS community?
We don’t appear to have gotten many practical proposals for how to make AIS more multidisciplinary, but there were a number of specific disciplines mentioned in the free responses, including cognitive psychology, neuroscience, game theory, behavioral science, ethics/law/sociology, and philosophy (epistemology was specifically brought up across multiple respondents). One respondent wrote, “AI alignment is dominated by computer scientists who don’t know much about human nature, and could benefit from more behavioral science expertise and game theory,” which I think captures the sentiment of many of the related responses most succinctly (however accurate this statement actually is!). Ultimately, encouraging and funding research at the intersection of these underexplored areas and alignment is likely the only thing that will actually lead to a more multidisciplinary research environment.
Did EAs consider AIS exclusively a longtermist cause area, or did they anticipate near-term catastrophic risk from AGI?
Unfortunately, I don’t think we asked the EA sample about AIS in a way that would allow us to answer this question using the data we have. This would be a really interesting follow-up direction. I will paste in below the ground truth distribution of EAs’ views on the relative promise of these approaches as additional context (eg, we see that the ‘AI risk’ and ‘Existential risk (general)’ distributions have very similar shapes), but I don’t think we can confidently say much about whether these risks were being conceptualized as short- or long-term.
It’s also important to highlight that in the alignment sample (from the other survey), researchers generally indicate that they do not think we’re going to get AGI in the next five years. Again, this doesn’t clarify if they think there are x-risks that could emerge in the nearer term from less-general-but-still-very-advanced AI, but it does provide an additional datapoint that if we are considering AI x-risks to be largely mediated by the advent of AGI, alignment researchers don’t seem to expect this as a whole in the very short term:
What links here?
- Cameron Berg's comment on Key takeaways from our EA and alignment research surveys by Cameron Berg (EA Forum; 6 May 2024 19:07 UTC; 2 points)

Cameron Berg 4 May 2024 15:45 UTC
7 points
1
in reply to: Josh Jacobson’s comment on: Key takeaways from our EA and alignment research surveys
I expect this to generally be a more junior group, often not fully employed in these roles, with eg the average age and funding level of the orgs that are being led particularly low (and some of the orgs being more informal).
Here is the full list of the alignment orgs who had at least one researcher complete the survey (and who also elected to share what org they are working for): OpenAI, Meta, Anthropic, FHI, CMU, Redwood Research, Dalhousie University, AI Safety Camp, Astera Institute, Atlas Computing Institute, Model Evaluation and Threat Research (METR, formerly ARC Evals), Apart Research, Astra Fellowship, AI Standards Lab, Confirm Solutions Inc., PAISRI, MATS, FOCAL, EffiSciences, FAR AI, aintelope, Constellation, Causal Incentives Working Group, Formalizing Boundaries, AISC.
~80% of the alignment sample is currently receiving funding of some form to pursue their work, and ~75% have been doing this work for >1 year. Seems to me like this is basically the population we were intending to sample.
One additional factor for my abandoning it was that I couldn’t imagine it drawing a useful response population anyway; the sample mentioned above is a significant surprise to me (even with my skepticism around the makeup of that population). Beyond the reasons I already described, I felt that it being done by a for-profit org that is a newcomer and probably largely unknown would dissuade a lot of people from responding (and/or providing fully candid answers to some questions).
Your expectation while taking the survey about whether we were going to be able to get a good sample does not say much about whether we did end up getting a good sample. Things that better tell us whether or not we got a good sample are, eg, the quality/distribution of the represented orgs and the quantity of actively-funded technical alignment researchers (both described above).
All in all, I expect that the respondent population skews heavily toward those who place a lower value on their time and are less involved.
Note that the survey took people ~15 minutes to complete and resulted in a $40 donation being made to a high-impact organization, which puts our valuation of an hour of their time at ~$160 (roughly equivalent to the hourly rate of someone who makes ~$330k annually). Assuming this population would generally donate a portion of their income to high-impact charities/organizations by default, taking the survey actually seems to probably have been worth everyone’s time in terms of EV.

Cameron Berg 3 May 2024 20:11 UTC
7 points
0
in reply to: Chris_Leong’s comment on: Key takeaways from our EA and alignment research surveys
There’s a lot of overlap between alignment researchers and the EA community, so I’m wondering how that was handled.
Agree that there is inherent/unavoidable overlap. As noted in the post, we were generally cautious about excluding participants from either sample for reasons you mention and also found that the key results we present here are robust to these kinds of changes in the filtration of either dataset (you can see and explore this for yourself here).
With this being said, we did ask in both the EA and the alignment survey to indicate the extent to which they are involved in alignment—note the significance of the difference here:
From alignment survey:
From EA survey:
This question/result serves both as a good filtering criterion for cleanly separating out EAs from alignment researchers and also gives a pretty strong evidence that we are drawing on completely different samples across these surveys (likely because we sourced the data for each survey through completely distinct channels).
Regarding the support for various cause areas, I’m pretty sure that you’ll find the support for AI Safety/Long-Termism/X-risk is higher among those most involved in EA than among those least involved. Part of this may be because of the number of jobs available in this cause area.
Interesting—I just tried to test this. It is a bit hard to find a variable in the EA dataset that would cleanly correspond to higher vs. lower overall involvement, but we can filter by number of years one has been involved involved in EA, and there is no level-of-experience threshold I could find where there are statistically significant differences in EAs’ views on how promising AI x-risk is. (Note that years of experience in EA may not be the best proxy for what you are asking, but is likely the best we’ve got to tackle this specific question.)
Blue is >1 year experience, red is <1 year experience:
Blue is >2 years experience, red is <2 years experience:

Key takeaways from our EA and alignment research surveys

Cameron Berg, Judd Rosenblatt, florin_pop and AE Studio

3 May 2024 18:10 UTC

94 points

10 comments21 min readLW link

Cameron Berg 27 Mar 2024 0:13 UTC
3 points
0
in reply to: Joseph Miller’s comment on: AE Studio @ SXSW: We need more AI consciousness research (and further resources)
Thanks for the comment!
Consciousness does not have a commonly agreed upon definition. The question of whether an AI is conscious cannot be answered until you choose a precise definition of consciousness, at which point the question falls out of the realm of philosophy into standard science.
Agree. Also happen to think that there are basic conflations/confusions that tend to go on in these conversations (eg, self-consciousness vs. consciousness) that make the task of defining what we mean by consciousness more arduous and confusing than it likely needs to be (which isn’t to say that defining consciousness is easy). I would analogize consciousness to intelligence in terms of its difficulty to nail down precisely, but I don’t think there is anything philosophically special about consciousness that inherently eludes modeling.
is there some secret sauce that makes the algorithm [that underpins consciousness] special and different from all currently known algorithms, such that if we understood it we would suddenly feel enlightened? I doubt it. I expect we will just find a big pile of heuristics and optimization procedures that are fundamentally familiar to computer science.
Largely agree with this too—it very well may be the case (as seems now to be obviously true of intelligence) that there is no one ‘master’ algorithm that underlies the whole phenomenon, but rather as you say, a big pile of smaller procedures, heuristics, etc. So be it—we definitely want to better understand (for reasons explained in the post) what set of potentially-individually-unimpressive algorithms, when run in concert, give you system that is conscious.
So, to your point, there is not necessarily any one ‘deep secret’ to uncover that will crack the mystery (though we think, eg, Graziano’s AST might be a strong candidate solution for at least part of this mystery), but I would still think that (1) it is worthwhile to attempt to model the functional role of consciousness, and that (2) whether we actually have better or worse models of consciousness matters tremendously.

AE Studio @ SXSW: We need more AI consciousness research (and further resources)

AE Studio, Cameron Berg, Judd Rosenblatt, phgubbins and Diogo de Lucena

26 Mar 2024 20:59 UTC

66 points

7 comments3 min readLW link

Cameron Berg 23 Feb 2024 18:10 UTC
1 point
0
in reply to: Kajus’s comment on: Survey for alignment researchers!
There will be places on the form to indicate exactly this sort of information :) we’d encourage anyone who is associated with alignment to take the survey.

Cameron Berg 9 Feb 2024 14:08 UTC
1 point
0
in reply to: Linda Linsefors’s comment on: Survey for alignment researchers!
Thanks for taking the survey! When we estimated how long it would take, we didn’t count how long it would take to answer the optional open-ended questions, because we figured that those who are sufficiently time constrained that they would actually care a lot about the time estimate would not spend the additional time writing in responses.
In general, the survey does seem to take respondents approximately 10-20 minutes to complete. As noted in another comment below,
this still works out to donating $120-240/researcher-hour to high-impact alignment orgs (plus whatever the value is of the comparison of one’s individual results to that of community), which hopefully is worth the time investment :)

Cameron Berg 8 Feb 2024 15:25 UTC
1 point
0
in reply to: Michael Tontchev’s comment on: Survey for alignment researchers!
Ideally within the next month or so. There are a few other control populations still left to sample, as well as actually doing all of the analysis.

Cameron Berg 7 Feb 2024 14:10 UTC
2 points
0
in reply to: Esben Kran’s comment on: Survey for alignment researchers!
Thanks for sharing this! Will definitely take a look at this in the context of what we find and see if we are capturing any similar sentiment.

Survey for alignment researchers!

Cameron Berg, Judd Rosenblatt and AE Studio

2 Feb 2024 20:41 UTC

71 points

11 comments1 min readLW link

Cameron Berg 19 Dec 2023 18:10 UTC
2 points
0
in reply to: Roman Leventov’s comment on: The ‘Neglected Approaches’ Approach: AE Studio’s Alignment Agenda
Thanks for calling this out—we’re definitely open to discussing potential opportunities for collaboration/engaging with the platform!

Cameron Berg 19 Dec 2023 18:07 UTC
3 points
0
in reply to: Roman Leventov’s comment on: The ‘Neglected Approaches’ Approach: AE Studio’s Alignment Agenda
It’s a great point that the broader social and economic implications of BCI extend beyond the control of any single company, AE no doubt included. Still, while bandwidth and noisiness of the tech are potentially orthogonal to one’s intentions, companies with unambiguous humanity-forward missions (like AE) are far more likely to actually care about the societal implications, and therefore, to build BCI that attempts to address these concerns at the ground level.
In general, we expect the by-default path to powerful BCI (i.e., one where we are completely uninvolved) to be negative/rife with s-risks/significant invasions of privacy and autonomy, etc, which is why we are actively working to nudge the developmental trajectory of BCI in a more positive direction—i.e., one where the only major incentive is build the most human-flourishing-conducive BCI tech we possibly can.

Cameron Berg 19 Dec 2023 16:31 UTC
8 points
4
in reply to: Roman Leventov’s comment on: The ‘Neglected Approaches’ Approach: AE Studio’s Alignment Agenda
With respect to the RLNF idea, we are definitely very sympathetic to wireheading concerns. We think that approach is promising if we are able to obtain better reward signals given all of the sub-symbolic information that neural signals can offer in order to better understand human intent, but as you correctly pointed out that can be used to better trick the human evaluator as well. We think this already happens to a lesser extent and we expect that both current methods and future ones have to account for this particular risk.
More generally, we strongly agree that building out BCI is like a tightrope walk. Our original theory of change explicitly focuses on this: in expectation, BCI is not going to be built safely by giant tech companies of the world, largely given short-term profit-related incentives—which is why we want to build it ourselves as a bootstrapped company whose revenue has come from things other than BCI. Accordingly, we can focus on walking this BCI developmental tightrope safely and for the benefit of humanity without worrying if we profit from this work.
We do call some of these concerns out in the post, eg:
We also recognize that many of these proposals have a double-edged sword quality that requires extremely careful consideration—e.g., building BCI that makes humans more competent could also make bad actors more competent, give AI systems manipulation-conducive information about the processes of our cognition that we don’t even know, and so on. We take these risks very seriously and think that any well-defined alignment agenda must also put forward a convincing plan for avoiding them (with full knowledge of the fact that if they can’t be avoided, they are not viable directions.)
Overall—in spite of the double-edged nature of alignment work potentially facilitating capabilities breakthroughs—we think it is critical to avoid base rate neglect in acknowledging how unbelievably aggressively people (who are generally alignment-ambivalent) are now pushing forward capabilities work. Against this base rate, we suspect our contributions to inadvertently pushing forward capabilities will be relatively negligible. This does not imply that we shouldn’t be extremely cautious, have rigorous info/exfohazard standards, think carefully about unintended consequences, etc—it just means that we want to be pragmatic about the fact that we can help solve alignment while being reasonably confident that the overall expected value of this work will outweigh the overall expected harm (again, especially given the incredibly high, already-happening background rate of alignment-ambivalent capabilities progress).

Cameron Berg 19 Dec 2023 16:11 UTC
6 points
2
in reply to: Roman Leventov’s comment on: The ‘Neglected Approaches’ Approach: AE Studio’s Alignment Agenda
Thanks for your comment! I think we can simultaneously (1) strongly agree with the premise that in order for AGI to go well (or at the very least, not catastrophically poorly), society needs to adopt a multidisciplinary, multipolar approach that takes into account broader civilizational risks and pitfalls, and (2) have fairly high confidence that within the space of all possible useful things to do to within this broader scope, the list of neglected approaches we present above does a reasonable job of documenting some of the places where we specifically think AE has comparative advantage/the potential to strongly contribute over relatively short time horizons. So, to directly answer:
Is this a deliberate choice of narrowing your direct, object-level technical work to alignment (because you think this where the predispositions of your team are?), or a disagreement with more systemic views on “what we should work on to reduce the AI risks?”
It is something far more like a deliberate choice than a systemic disagreement. We are also very interested and open to broader models of how control theory, game theory, information security, etc have consequences for alignment (e.g., see ideas 6 and 10 for examples of nontechnical things we think we could likely help with). To the degree that these sorts of things can be thought of further neglected approaches, we may indeed agree that they are worthwhile for us to consider pursuing or at least help facilitate others’ pursuits—with the comparative advantage caveat stated previously.
What links here?
- Roman Leventov's comment on The ‘Neglected Approaches’ Approach: AE Studio’s Alignment Agenda by Cameron Berg (19 Dec 2023 17:33 UTC; 2 points)

The ‘Neglected Approaches’ Approach: AE Studio’s Alignment Agenda

Cameron Berg, Judd Rosenblatt, AE Studio and Marc Carauleanu

18 Dec 2023 20:35 UTC

160 points

20 comments12 min readLW link

Computational signatures of psychopathy

Cameron Berg19 Dec 2022 17:01 UTC

29 points

3 comments20 min readLW link

Cameron Berg 15 Dec 2022 23:09 UTC
7 points
3
on: Consider working more hours and taking more stimulants
I’m definitely sympathetic to the general argument here as I understand it: something like, it is better to be more productive when what you’re working towards has high EV, and stimulants are one underutilized strategy for being more productive. But I have concerns about the generality of your conclusion: (1) blanket-endorsing or otherwise equating the advantages and disadvantages of all of the things on the y-axis of that plot is painting with too broad a brush. They vary, eg, in addictive potential, demonstrated medical benefit, cost of maintenance, etc. (2) Relatedly, some of these drugs (e.g., Adderall) alter the dopaminergic calibration in the brain, which can lead to significant personality/epistemology changes, typically as a result of modulating people’s risk-taking/reward-seeking trade-offs. Similar dopamine agonist drugs used to treat Parkinson’s led to pathological gambling behaviors in patients who took it. There is an argument to be made for at least some subset of these substances that the trouble induced by these kinds of personality changes may plausibly outweigh the productivity gains of taking the drugs in the first place.

Cameron Berg 24 Oct 2022 22:17 UTC
10 points
3
in reply to: paulfchristiano’s comment on: AI researchers announce NeuroAI agenda
27 people holding the view is not a counterexample to the claim that it is becoming less popular.
Still feels worthwhile to emphasize that some of these 27 people are, eg, Chief AI Scientist at Meta, co-director of CIFAR, DeepMind staff researchers, etc.
These people are major decision-makers in some of the world’s leading and most well-resourced AI labs, so we should probably pay attention to where they think AI research should go in the short-term—they are among the people who could actually take it there.
See also this survey of NLP
I assume this is the chart you’re referring to. I take your point that you see these numbers as increasing or decreasing (despite that where they actually are in an absolute sense seems harmonious with believing that brain-based AGI is entirely possible), but it’s likely that these increases or decreases are themselves risky statistics to extrapolate. These sorts of trends could easily asymptote or reverse given volatile field dynamics. For instance, if we linearly extrapolate from the two stats you provided (5% believe scaling could solve everything in 2018; 17% believe it in 2022), this would predict, eg, 56% of NLP researchers in 2035 would believe scaling could solve everything. Do you actually think something in this ballpark is likely?
Did the paper say that NeuroAI is looking increasingly likely?
I was considering the paper itself as evidence that NeuroAI is looking increasingly likely.
When people who run many of the world’s leading AI labs say they want to devote resources to building NeuroAI in the hopes of getting AGI, I am considering that as a pretty good reason to believe that brain-like AGI is more probable than I thought it was before reading the paper. Do you think this is a mistake?
Certainly, to your point, signaling an intention to try X is not the same as successfully doing X, especially in the world of AI research. But again, if anyone were to be able to push AI research in the direction of being brain-based, would it not be these sorts of labs?
To be clear, I do not personally think that prosaic AGI and brain-based AGI are necessarily mutually exclusive—eg, brains may be performing computations that we ultimately realize are some emergent product of prosaic AI methods that already basically exist. I do think that the publication of this paper gives us good reason to believe that brain-like AGI is more probable than we might have thought it was, eg, two weeks ago.

Cameron Berg

There Should Be More Align­ment-Driven Startups

Key take­aways from our EA and al­ign­ment re­search sur­veys

AE Stu­dio @ SXSW: We need more AI con­scious­ness re­search (and fur­ther re­sources)

Sur­vey for al­ign­ment re­searchers!

The ‘Ne­glected Ap­proaches’ Ap­proach: AE Stu­dio’s Align­ment Agenda

Com­pu­ta­tional sig­na­tures of psychopathy