Simon Skade
I did (mostly non-prosaic) alignment research between Feb 2022 and Aug 2025. (Won $10k in the ELK contest, participated in MLAB and SERI MATS 3.0 & 3.1, then independent research. I worked a bit on ontology identification and then on an ambitious attempt to better understand minds to figure out how to create more understandable and pointable AIs. I started with agent foundations but then developed a more sciency agenda where I also studied concrete observations from language/linguistics, a little bit pychology and neuroscience, and from tracking my thoughts on problems I solved (aka a good kind of introspection).)
I spent the last 8 months exploring AI governance and public and political advocacy about AI risks. (As of June 2026, I am now considering what to do next.)
I’m also into rationality/self-improvement.
Currently based in Germany.
I think it’s important that we distinguish between slopiness from AIs not having been taught to try hard with sensible approaches (elicitation) and slopiness because the AI capability profile is shaped in a way that AIs are just not very good at some tasks. Though I agree it’s a spectrum.
I think it’s probably not the case that the profile of potential AI capabilities (aka the profile if all capabilities were well elicited) is very similar to that of humans. In particular, I expect AIs to be e.g. differentially quite bad at e.g. MIRI-like conceptual alignment research where you need to grapple very deeply with a phenomenon or puzzle and form good new concepts/ontologies/frameworks to get a handle on that. Or like deriving general relativity from as little evidence as Einstein did. I.e. even if we tried hard and had more useful data here we might have a hard time to make AIs good at that. (E.g. maybe current AIs may lack some cognitive machinery part that humans have for forming natural ontologies. Or for noticing confusions / seeing what evidence is suprising given the world model and using that to find where beliefs may be wrong.)
So I guess actually like 3 levels on a spectrum:
actual elicited capabilities
capability profile (what could be elicited from an AI model with good post-training (aka assuming you have good datasets/tests))
architecture potential profile
(I realize this is only tangential to your comment, but I’ve been reading a bunch of stuff from Ryan lately and I felt this distinction wasn’t made properly but there’s no clear location where to best comment it and I just felt an impulse now.)
(Or like, Ryan mentions both 1 and 2, but not sure if he thinks there’s a meaningful difference on 3 to humans. Personally 3 does make me more pressimistic about relative speedups of alignment research vs capability research. But tbc it could be that 3 is still not that important if we find kinds of safety research that don’t need that much of forming deep novel ontologies or so. But it could also be important.)