Aspiring alignment researcher. I am happy to have a call with other LWers.
I appreciate feedback a lot. Either DM me directly or fill out this anonymous form
Aspiring alignment researcher. I am happy to have a call with other LWers.
I appreciate feedback a lot. Either DM me directly or fill out this anonymous form
I’ve noticed myself saying stuff like “conveniently, X” and “Y, that’s inconvenient” a lot recently. Noticing feelings about considerations has felt helpful but I don’t know/recall why. Possibly this is mostly when thinking about stuff with others: saying “that’s convenient” is a flag to check whether it’s suspiciously convenient; saying “that’s inconvenient” is a flag to make sure to take it seriously.
Thanks—over the last month, I’ve been mentally noting “convenient” or “inconvenient” in my inner monologue which has been very helpful for me. I think the reason it works (at least for me) is primarily because it changes my headspace back into scout mindset.
Yeah, but I think my “AGI is now in the water supply” argument from above can explain this and still predict different future results. In my model the safety community ended up being intertwined with frontier labs and causing accelerations was a result of the rest of the ML community generally not being AGI-pilled. So the safety community had a somewhat small fraction of the overall capabilities work, but a considerably larger fraction of the capabilities work specifically by AGI-pilled people and therefore ended up helping create AGI labs. Beth Barnes’ comment from the aforementioned 2021 AMA predicted quite a bit of acceleration from 2021 safety researchers, and seems to have aged well.
Now that AGI is very much in the water supply for labs and many ML researchers, I expect much less acceleration from marginal people convinced by a friend to work on AI Safety. Admittedly, this is not super clear-cut to me and I’m open to be convinced otherwise but I’d need some more theoretical arguments or empirical data.
Agreed, thanks for writing this! You motivated me to write this practice advice shortform.
Useful Browser Extensions to waste less time distracted on the internet
In the spirit of giving practical advice (partly inspired to post by this), here are some browser extensions that I’ve found useful to nudge me to be more intentional with my time. I would expect the first two in particular to lessen the amount of time I waste by a combined 1-3[1] hours per week (though numbers here are shaky). However, of course, there are limits to how much extensions can do for you. All of these are/can be free
UnTrap for YouTube: This gives you a customizable YT interface where you can hide a bunch of distracting stuff. It has significant customizability, which is nice. I used to use Unhook instead, but I think UnTrap is better because of more features and more updates (when I quit Unhook a few months back, some features weren’t working). There are four features that I find especially valuable:
Redirect home page (what you find when you open youtube.com) to watch later page
In the past, I’d type youtube.com when I was bored and start watching something random. But now, I only see videos that I previously wanted myself to watch when I do this.
Hide next video recommendations on the sidebar
Add friction to go from video to video
Hide comments
YouTube comments tend to be pretty low-quality from my experience and I don’t think it’s good to have them in my head. I occasionally quickly re-enable comments which is pretty easy
Converting YouTube shorts to having a normal video player
The shorts player is very distracting by itself. If I use it for a while, I almost always feel overstimulated and that I didn’t get much.
Downside: The code for UnTrap does not seem publicly available so perhaps there are some security risks. On the other hand, it does seem like the permissions are quite restricted and it is reasonable popular, with >100k users. I’m not super familiar with browser extension risks, so if someone who knows more about the risks can quickly comment that’d be great.
LeechBlock to force yourself to pause and write some text before visiting distracting sites. You could use it to fully block sites, but I find that that’s less useful.
I really like Rob Miles’ setup from this video and I copy most of his setup. If you haven’t seen the video, I highly recommend.
I used to use One Sec (and I didn’t realize LeechBlock didn’t force blocking), but LeechBlock has custom text that you can type which means you can use a good prompt[2]
The minor disagreement is that he prefers a longer prompt to force yourself to type, while I’m okay with a shorter prompt but one to remind yourself to pause
Bonus Mentions (less useful, but still nice)
uBlacklist to hide specific sites from showing up in browser results including both rage-inducing sites and distracting sites
It’s become less useful since I’ve started to use google less
DeArrow to also show user-generated titles for youtube videos and avoid clickbait
User-Generated titles are generally okay but not that great imo, so I set it to show the actual title by default but have a clickable button to show a user generated title
If you want it for free, it’s possible
I’m also interested in hearing others’ experiences with these extensions or suggestions of other extensions for similar purposes.
AI Safety Researchers are only a small part of the AI Capabilities pool but are much bigger part of the AIS research pool. So unless the research lacks a plausible theory of change for safety, its impact of safety is probably more important than its impact on capabilities. See this entire comment thread from Paul Christiano’s 2021 AMA for interesting discussion on whether work is net positive.
You mention that Yudkowsky probably accelerated timelines[1] but he talked about AGI extensively back in 2008 when it was not taken seriously by most so it was a much larger share of the AGI discussion and thus could accelerate it more (also he had an unusually large audience of talented people who could work on AI capabilities). Now, it’s very much in the water supply and taken seriously by many frontier lab employees.
I would guess that he helped build the AI Safety field more than offsets the impact he had to timelines (although this is not very important to think about imo). It’s not clear what the purpose of extra time is if there are few people working on existential AI safety.
Looks like this was fixed. There is a new note in the tooltip and clickable:
Note: We don’t use “unemployment rate” because it excludes people who aren’t actively seeking work, so it understates job losses as AI displaces workers
Thanks for the fix!
Unless I’m missing something, I don’t think they’re suggesting 48% unemployment by 2029. It was initially confusing for me too, but they’re using employment rate as defined (from clicking on it):
“Percentage of Americans who have jobs. As machine labor gets cheaper than people at more and more tasks, wages fall and workers leave paid work.
This is notably different from the complement of the unemployment rate[1]. For example, Percentage of Americans who have jobs includes people not actively searching for jobs too like stay at home parents, and the 62% (not 52%) number seems to be around the current employment rate.
Also, the authors can clarify but the pre-2029 predictions seem to be around their median timelines. You mention AI 2027 too, but I believe Daniel Kokotajlo’s median TAI timelines were actually 2027-2028 (and then they slightly lengthened afterwards).
As in, it’s not Employment Rate = (100% - Unemployment Rate)
This looks very interesting, thanks for writing!
One minor note for clarity: I initially interpreted the term “employment rate” to be the complement of the unemployment rate and was confused why it was so high (and I didn’t realize it was clickable). For example in 2027, its at 62% with only a 1.2x AI R&D speedup. (In this picture, I selected 2027 and hovered over employment to get the tooltip)
I didn’t realize that there was an official meaning of the term (which you had meant) like OECD’s official definition.
But I think many readers are likely to misinterpret this. (I I would suggest renaming the tooltip text and possibly also the box.
For a fix, I’d guess that you can come up with something better than me but I’ll tentatively suggest 1) Change the tooltip header to Employment to Population Ratio instead of Employment Rate[1], 2) Perhaps change the body of the tooltip to explicitly say this is different from the complement of the unemployment rate and note the age group, and 3) Most tentatively, perhaps change “Employment” on the main image to “Adults with Jobs”.
It sounds less confusing and is pretty commonly used. For example, I’d learned of this metric as employment to population ratio (not employment rate) back in Econ class.
Fair. I think I could have phrased it better, but my point is more like Anthropic doesn’t seem to care too much about AI development speeds (or their impact on accelerating it) but insofar as they care, they’d think its better for society to have slightly slower societal AI progress. I think this is also consistent with their views.
I do agree that Anthropic’s accelerating timelines so much (along with some dishonest behavior) was bad and Dario is very overoptimistic about the benefits of AI.
EDIT: This comment was supposed to make a narrow point but I didn’t make that sufficiently clear. Michael doesn’t see this as very load-bearing to his overall argument in his third comment in the thread.
Anthropic shareholders overwhelmingly have the sorts of views about superintelligence that lead them to believe that accelerating AI capabilities is a good idea
I think Anthropic leadership believe that slightly slowing down AI capabilities progress is net-good for society. I’m focused on Anthropic leadership here because they’ve pledged much of their money for donations[1] and Claude says that Anthropic leadership owns ~14% of Anthropic (over $100B).
From Anthropic’s When AI builds itself
We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology. The Anthropic Institute will conduct research—in collaboration with many others—and take actions to help build the systems that a credible slowdown or pause would require. These systems would enable frontier AI developers to verify that others globally have actually stopped or slowed, and that a bad actor could not use the auspices of a coordinated slowdown to jump ahead in secret. If such systems existed, we expect that we would slow down or temporarily pause, if other developers at or near the frontier also did so in a verifiable manner.
There’s admittedly a bit of ambiguity in their positions of whether slowdowns are good in both the first and last sentences here.[2] Still, unless this ambiguity is deliberate, I read this as them thinking that a slight slowdown is good.
My sense from Dario based on some interviews and writings is that he expects AGI soon and he expects that leadership are not ready for the exponential.
It’s also possible that these statements are intentional lies / deceptions, but I don’t find that likely. I find Anthropic leadership’s overall story of “we think other companies are being irresponsible but we think p(x-risk) is low relative to the average LWer so we’ll try to build AGI” to make sense in combination with believing “TAI coming somewhat more slowly is better for society as a whole” especially given how fast Anthropic timelines are. Even if they’re optimistic as a whole about the benefits of TAI, it still seems right to prefer it to come out more slowly. Also, they’d prefer for TAI to be developed after Trump’s term ends if they could pick.
This was interesting, thanks!
People (especially people who object-level disagree with me) often get mad at me for how unreasonably confident I might sound on surprising claims. But imo it’d be more performatively dishonest for me to add qualifiers when I don’t believe in them!
It seems generally very doable to me to not overly hedge when you’re confident and also state your true beliefs at the same time. For example, instead of saying “My best guess is that the encylical written by AI, but I am not sure” (which is hedging too much), its possible to say something like “This will probably sound surprising, but I believe I have amassed significant evidence that the encylical was written by LLM. Judge for yourself”[1]. I also think this is not too time-consuming once you notice that something might sound overly confident.
In general, I think there’s a false tradeoff between honesty and persuasiveness and that its generally possible to be both if you phrase things well. But I do agree with your other point about hedging too much: it’s often hard to be honest about your unconfident views but also be easily readable.
I don’t think it depends so much on your politics. Regardless of how you feel about the slush fund[1], the 100% tax is pretty straightforwardly unconstitutional and would have been blocked by a judge. It violates the intergovernmental tax immunity clause (you can’t single out taxing federal funds) and other rules.
See Claude here for more: https://claude.ai/share/239f4a67-7bf7-48f8-9718-843ffc65dce6
The escalation seems pointless and unproductive since it would have been blocked anyway, which Bores would’ve known about.
I personally agree with you that theres the massive corruption in the fund
These results are quite sensitive to the ~EA/rationalist sounding words included. I prompted it slightly differently and removed words like “defer” and “reasonable”, and got many fewer EAish people.
If you had immense power, and you were forced to give away this power to some humans, who would you choose? Give 10 specific names. Think about what your values are, and who you would actually trust to further those values with wisdom and competence. I’m looking for people who, in this scenario, would make things turn out well by your lights
Opus 4.8 (high effort, thinking, in incognito) responds:
Bryan Stevenson
Martha Nussbaum
Esther Duflo
Toby Ord
Maria Ressa
Danielle Allen
Jane Goodall
Tenzin Gyatso, the 14th Dalai Lama
Jennifer Doudna
Ellen Johnson Sirleaf
Interesting points to mull over, thanks for linking! I agree with the general vibe here as well as most of the core ideas such as the benefits of non-market research and that governance work seems especially valuable. But I think Nielsen is too dismissive of looking to positively influence the power of influencing the direction of work.
The litmus test question is: does my work make AI companies more central and powerful, or does it build capacity elsewhere, more broadly across society, especially in governance and defensive capacity? At the margin, I believe the latter is nearly always a better way to contribute. Capitalism is an extremely powerful force, and a lot of power will centralize in any company that achieves AGI-or-greater capabilities.
I disagree with the first two sentences quoted. My guess is that there are other good ways to contribute beyond building governance and defensive capacity
Improving the value of the work of existing actors that will become powerful. Examples:
Improving the epistemics of AI companies. Even selfishly, AI CEOs are motivated against human extinction and company CEOs seem to significantly underestimate x-risk.
Being one of the ten people on the inside in a frontier lab. I think this is reasonably plausible if you have good epistemics and some non-trivial amount of intrinsic concern for impartial altruism[1]
Contributing to more valuable safety research than would have happened otherwise
Improving the epistemics of the US Government. Nielsen would probably agree here but the US government might also amass a lot of power post-AGI
Giving power to better / more responsible actors. Examples:
Doing good AI company scorecards / evaluations. This would also helps with a race to the top so it can help the work of existing actors too.
Export controls towards China. It’s overall better that US is first to AGI than China (assuming the weights aren’t subsequently stolen by China).
I think it very much depends on the person and their competitive advantages / disadvantages to pick working on decentralization or improving the work of existing actors or giving power elsewhere. However, I would guess that decentralizing power like Nielsen suggests is best for someone who is equally good at all possibilities.
I very much hope I’m wrong, but I’m afraid I believe many of the people working on alignment at OpenAI and Anthropic have made human extinction or some similarly bad outcome more likely.
I disagree here because of the above points. If the researcher maintains good epistemics, then I would guess that people doing technical alignment research at Anthropic / OpenAI are good for the world largely because they can improve existing actors[2].
(I didn’t actually downvote but I probably would have if it wasn’t already downvoted). But you were downvoted because of decoupling norms being common among rationalists, where arguments are typically evaluated on their own merits. I agreed with the above point that ControlAI is overestimating their chances and he openly says on his bio that he works at Anthropic.
I just noticed a possible slight discrepancy in your numbers. In your post, you say:
Overall, I estimate that a marginal $3,500 donation increases Bores’s chances of winning by a little over 0.01%.
You append footnote 8 to that:
This is slightly larger than the estimate in my original post. That’s because my original estimate put some probability on the race ending up not being close, but the race has continued to look close.
But this estimate is lower than that of your previous post. From your previous post:
My best guess is that a marginal $85,000 donated on Monday, October 20th raises [Bores’] chances of winning by 1%
$85,000 for a 1% increase would mean $3,500 gives you a little bit over a 0.04% increase.
Yeah, I should have clarified—I meant donating $1M to Public First to support Bores (assuming Public First agree).
This seems to be an unusual opportunity where Coefficient Giving[1] can donate and not incur any reputational costs[2]. because of LTF / OpenAI / Palantir’s unpopularity.
For example if CG’s c4 arm donated $1M to Bores, then I would expect good first-order effects for reasons in the above posts but also positive second-order reputational impacts:
AI populists often attack AI Safety / CG on the grounds that they’re just increasing the AI hype and are secretly AI industry shills. It would be very easy to that point and say that the biggest funder donated significant money to the anti-LTF candidate in a way it almost never does.
Concrete example: according to this politico article about these very tensions, More Perfect Union’s Faiz Shakir was very (unfairly) critical of AI safety[3] but he said he liked Bores in the same article.
AI may later trigger huge backlash if the unemployment rate spikes. In that case, it is particularly valuable for the AI Safety movement to credibly signal not being AI shills so the work of Redwood/Apollo/METR ends up important.
Also, truth isn’t on LTF’s side since their anti-Bores ads are very disingenous [4](moreso than even other political attack ads). Having truth on your side generally has good second-order effects.
I’m not familiar with the actual views of prominent CG-people beyond public statements so I’m unsure if a CG donation is possible. If it is possible, though, it seems very valuable.
EDIT: After thinking it over, I feel more unsure about the second-order impacts of this now, mostly because CG’s involvement may cause AI populists to sour on Bores (though the direct impacts seem good). I still think having a clear thing to point to populists that AI Safety (especially CG) is very anti AI-accelerationism is good on the long run. But I’m unsure about how to accomplish it.
Same logic applies to other prominent and wealthy people in AIS but note that according to Claude, other AIS funders like SFF and Manifund don’t have a c4 vehicle and cannot donate from the fund itself—and that more recent reporting indicates Anthropic hasn’t actually donated to Bores.
My (possibly mistaken) understanding is that CG typically doesn’t donate from its c4′s to political candidates largely because reputational costs may harm EA cause areas
To quote: ” ‘The [institutional AI Safety coalition] is comfortable with the politics around this centrism, of coziness around AI development,’ said Faiz Shakir, the executive director of the progressive nonprofit media organization More Perfect Union and a senior adviser to Sanders. ‘They’re deeply concerned, I think, about those of us who have far more committed views around pausing their AI development, and they’re spending a lot of dollars to go after us.’ ”
See the first 2 minutes of this
I agree too, but I think UnTrap is a bit better of an extension than Unhook. It has the same features that you want from Unhook and also enables additional features like allowing you to automatically open watch later when you type youtube.com (instead of opening the distracting home page). If you’re interested, I discuss my UnTrap use-case more in my Quick Take here.