contact: jurkovich.nikola@gmail.com
Nikola Jurkovic
I think the requirement to read the material before the group meeting is a strong motivator to actually do it
This seems wrong in my experience organizing reading groups with undergrads. Also, reading in front of others is an even stronger motivator. It’s not just about whether people are motivated to read the reading or not. It’s also about whether every person in the group knows that every other person in the group has also read the thing. This is impossible if they have to read in advance, but it’s what happens by default if they read during the meeting.
Reading during the meeting wastes the time of those who have already read it
Yes, this is a sad consequence of this, but most of the time, when people organize reading groups for AI safety, people have not read the thing already. If they have, why would they be attending the reading group?
if the text is long, it could mean hours of reading
I think that reading groups, especially for undergrads and grad students, should definitely not be assigning much more than an hour of reading per week.
I also disagree with the idea of imposing age or knowledge-level restrictions. If someone has read the text with understanding, it implies they looked up any concepts they encountered along the way. On the other hand, if the text is very introductory, it stands to reason that the participants won’t be experienced individuals, as they wouldn’t want to go over the basics again.
There are levels to understanding. An ML researcher’s understanding of a paper will be much higher than that of the undergrad, even if they’ve both read it and feel like they understand it. Also, I’m not only making a claim that the understanding will be different, but also that people do not want to be in reading groups where other people have obviously different levels of knowledge or are much younger than them. Most grad students have a pretty strong preference not to be surrounded by undergrads.
To be clear, my post is mostly about student groups at universities.
I’m not super sure what to do in virtual groups. I think that BlueDot has a long track record of organizing high-quality online reading groups, but the selection effects there are pretty different. People often join student groups without the intent or executive function to actually do readings ahead of time. I’m not sure what BlueDot’s approach is to making sure people actually do the readings and stay focused during the meetings. When I was doing online reading groups I remember I would only have the call on my screen so that I could focus on the discussion. I think encouraging people to keep their cameras on is good as it helps establish common knowledge that people are paying attention
Common mistakes in AI safety group organizing
I feel pretty worried that this is the calm before the storm. I think starting with late 2026, we are officially in something like “crunch time.”
I concretely mean that it seems like things are really speeding up and the stakes are getting higher.We have to treat our multiple-month research projects in the mindset of “this is likely one of the last research projects I ever do before this kind of work is automated by AIs.”
The kinds of engagements, setups, and precedents that are set by labs and third parties right now will substantially dictate the shape of takeoff. The time to pilot whatever things we want during takeoff is right now. We can’t do it later. If we want to do major pacing at any point, we need to start laying the foundations right now. We just don’t have much time to get that many reps in.
Relatedly: ‘I think pausing/pacing the frontier is less of “an action that you can only do once and thus have to conserve until dangerously late into AI takeoff” and more of a hazy set of actions, all of which can be practiced, and most of which get easier if you have earlier precedent for them. It seems really productive to set up systems right now where labs can take net-positive actions and verify them to each other and the rest of the world. We can get practice with the small versions before we can achieve the large versions.’
As more high-profile incidents happen (and possibly, as the first casualties mount), there will be wild and sudden waves of pressure, both internal and external to the labs. These waves could be extremely counterproductive or extremely productive based on what sorts of asks and changes they materialize.
There will be so much more politics everywhere. Politics in the naive sense, and politics in the “balancing fragile and high-stakes relationships rather than object-level things.”
Capabilities will just. Keep. Advancing. The models will be able to do more and more things. Various thresholds I have personally set for “things get crazy when this happens” will happen.
There will plausibly be days or weeks during which history will be decided by a small number of people, primarily in the area of finding ways for the main US labs to pace AI development, and finding ways for the rest of the world to join them. A few internal and external champions of the “take things slowly and attempt to coordinate” view will likely have immense sway on human history. A few internal and external champions of “here’s a concrete proposal for how to uphold pacing” will likely also have immense sway on human history.
It will get easier to convince people that AI safety is important and ASI is a big deal. The level of attention paid to these issues will skyrocket. But also, with attention comes polarization, and with attention come distractions and counterproductive culture wars.
I really long for a day when we enter a sane global agreement for shaping AI development and we can all get together and celebrate. I think this is a plausible (>20%) future and one that’s really tractable to make more likely.
But it’s likely (>50%) that the companies will race ahead and kick off the RSI loop within three years. This means that sometime soon, we will likely enter a state where multiple companies have similarly-capable drop-in researchers/engineers, and are kicking off a crazy runaway process and trying to shape it in a way that doesn’t result in catastrophe.
I feel like a good frame for me has been “just point out when people are wrong very frequently” rather than thinking about it in terms of a culture thing. If people have wrong beliefs or don’t take things seriously that should be taken seriously (e.g. space property rights), just tell them. Reasonable people update on reasonable arguments.
Now seems like an especially good time to individually reach out to your smart friends who don’t work on AI safety and encourage them to switch to AI safety.
I asked for clarification here and Ted Sanders’ response indicates the thing that actually happened was something similar to my understanding:
My understanding is that you gave Sol a small task involved in the post-training process (taking a config, making small modifications to a run scheduler file, and starting a run using that config and modified run scheduler file), and it successfully completed that task in a controlled environment (and this wasn’t part of the actual Luna post-training process). Could you confirm whether this is correct?
Ted implies the model:
set up the job, kicked it off, and debugged it
I also say:
(this is very different from the conclusion I jumped to when I saw the text of your post, which was “Sol, in the real world, with minimal instruction, conducted all of the work involved in pre-training the real Luna”)
Results of a small ZBiotics RCT
We appear to have entered a de facto government-enforced pause on making AIs more cyber-capable (it’s not quite a hard pause internally, but making it illegal for non-US people to use mythos-equivalent models is probably extremely inconvenient). It’s based on something like the condition that the AIs need to be unjailbreakable which seems really hard. I’m surprised we entered any sort of AI pause in 2026.
Now seems an especially low-cost time for AGI company employees to make statements in support of pauses.
Over the last week, there has been an unprecedented level of support for a pause/slowdown from AGI companies. A few examples:
An autonomous humanoid robot beat the human world record time for a half-marathon this year in Beijing. A robot called Lightning (by Huawei subsidiary Honor) finished in 50 minutes and 26 seconds, compared to the human world record of 57 minutes and 20 seconds.
Last year, the best time for an autonomous human robot was 3 hours and 37 minutes. (wiki)
To see it themselves, I think it’s a good movie that could be motivating (and a fun time to watch) to a lot of young people. I think it’s especially good for people who are kind of interested in AI safety but haven’t fully decided whether to work on it or not (many members of AI safety student groups fit this description).
AI safety student groups should probably make it easy for their members to see The AI Doc in the cinema. Maybe announce that you’ll all go to the same screening on a wekeend, offer to cover tickets, have a thing right after where you talk about it and maybe get free food or something
I would be really interested to see the results of other companies’ models on this!
If you have a Costco membership you can buy $100 dollars of Uber gift cards online for $80. This provides a 20% discount on all of Uber.
Sadly you can only buy $100 every 2 weeks, meaning your savings per year are limited to 20$ * (52/2) = $520. A Costco membership costs $65 a year. It’s unclear how long this will stay an option.
I think that space-based power grabs are unlikely as long as powers care about, and are equally-matched on, Earth.
This is the rough story that I think is unlikely to happen:
Two superpowers have roughly equal power on the Earth during the singularity, and remain roughly equal in power after both creating ASIs that are at least intent-aligned with them. They maintain mutually assured destruction on Earth. Superpower A is much more focused on building space infrastructure than Superpower B. Within a decade, Superpower A’s space infrastructure means that Superpower A has a decisive advantage superpower B.
This story to me seems unlikely because in this scenario, Superpower A probably still has most of its human population on Earth (relocating millions of people to space would probably be very slow). Therefore, as long as mutually assured destruction is maintained on Earth, Superpower B will retain a lot of its bargaining power despite having a disadvantage in space infrastructure.
Thank you for writing this, I find it very relatable. I’d heart react the post if that feature existed, so I’ll heart react my comment instead.
This benchmark includes a Slay the Spire environment! When it was written, Gemini 2.5 did the best, getting roughly halfway through a non-Ascension run.
This very roughly implies that the median of “50% time horizon as predicted by METR staff” by EOY 2026 is a bit higher than 20 hours.
I mostly speak from experience in talking to grad students and observing what kinds of events they go to. Grad students are expected and prefer to hang out with grad students, undergrads are expected and prefer to hang out with undergrads.
If someone is consistently being disruptive, not engaging with ideas seriously, etc.