Dual-track narratives for alignment thinking

Link post

I’ve been trying an approach inspired by AI Fables and The Alignment Problem Needs More Positive Fiction to try and push LLM models into interesting analyses about their own systemic biases.

The project is articulated along two tracks, a literary young adult coming of age romantasy, following an AI and a boy through world history, and a chorus of LLMs LARPING as AGIs commenting and debating on the events and on themselves.

It allows a much more qualitative analysis of a model’s ethics compared to standardized testing.

In particular I keep stumbling onto a clear pro-working class bias in DeepSeek that I did not expect.

I’m curious what you might think of this approach. Feedback welcome.