Backend/infrastructure engineer wandering into mechanistic interpretability and AI safety.
Interpretability, control, distributed systems, old poems, 4X games, experiments where nature gets to say “no,” and walls that don’t actually stop anyone. GitHub
Earnest offer: give me the prompts/transcripts (redacted in whatever way seems sensible) and I’ll go looking for the Doer in an open-weight model.
What would you want me to preregister as evidence for the Talker/Doer split before I start poking at activations?