Blaine

Karma: 174

Hi! I’m Blaine. I’m Research Communications Officer at an AGI company called Noeon Research, based in Japan. I run AI Safety 東京, a special interest group supporting Tokyo’s nascent AI safety scene. We run a yearly safety conference called TAIS.

https://aisafety.tokyo /

https://tais2024.cc/

https://noeon.ai/

https://linkedin.com/in/paperclipbadger

Tokyo AI Safety 2025: Call For Papers

Blaine21 Oct 2024 8:43 UTC

24 points

0 comments3 min readLW link

(www.tais2025.cc)

We ran an AI safety conference in Tokyo. It went really well. Come next year!

Blaine17 Jul 2024 6:55 UTC

46 points

1 comment6 min readLW link

Last call for submissions for TAIS 2024!

Blaine30 Jan 2024 12:08 UTC

4 points

0 comments1 min readLW link

(tais2024.cc)

Announcing TAIS 2024

Blaine6 Nov 2023 8:38 UTC

23 points

0 comments1 min readLW link

(tais2024.cc)

Blaine 13 Apr 2023 8:13 UTC
3 points
0
in reply to: Ulisse Mini’s comment on: Gradient Descent in Activation Space: a Tale of Two Papers
I’m not sure the tuned lens indicates that the model is doing iterative prediction; it shows that if for each layer in the model you train a linear classifier to predict the next token embedding from the activations, as you progress through the model the linear classifiers get more and more accurate. But that’s what we’d expect from any model, regardless of whether it was doing iterative prediction; each layer uses the features from the previous layer to calculate features that are more useful in the next layer. The inception network analysed in the distill.ai circuits thread starts by computing lines and gradients, then curves, then circles, then eyes, then faces, etc. Predicting the class from the presence of faces will be easier than from the presence of lines and gradients, so if you trained a tuned lens on inception v1 it would have the same pattern—lenses from later layers would have lower perplexity. I think to really show iterative prediction, you would have to be able to use the same lens for every layer; that would show that there is some consistent representation of the prediction being updated with each layer.

Here’s the relevant figure from the tuned lens—the transfer penalties for using a lens from one layer on another layer are small but meaningfully non-zero, and tend to increase the further away the layers are in the model. That they are small is suggestive that GPT might be doing something like iterative prediction, but the evidence isn’t compelling enough for my taste.

Blaine 12 Apr 2023 9:09 UTC
2 points
0
in reply to: Jon Garcia’s comment on: Gradient Descent in Activation Space: a Tale of Two Papers
Here’s a sketch of the predictive-coding-inspired model I think you propose:
The initial layer predicts token $i + 1$ from token $i$ for all tokens. The job of each “predictive coding” layer would be to read all the true tokens and predictions from the residual streams, find the error between the prediction and the ground truth, then make a uniform update to all tokens to correct those errors. As in the dual form of gradient descent, where updating all the training data to be closer to a random model also allows you to update a test output to be closer to the output of a trained model, updating all the predicted tokens uniformly also moves prediction $n + 1$ closer to the true token $n + 1$ . At the end, an output layer reads the prediction for $n + 1$ out of the latent stream of token $n$ .
This would be a cool way for language models to work:
- it puts next-token-prediction first and foremost, which is what we would expect for a model trained on next-token-prediction.
- it’s an intuitive framing for people familiar with making iterative updates to models / predictions
- it’s very interpretable, at each step we can read off the model’s current prediction from the latent stream of the final token (and because the architecture is horizontally homogenous, we can read off the model’s “predictions” for mid-sequence tokens too, though as you say they wouldn’t be quite the same as the predictions you would get for truncated sequences).
But we have no idea if GPT works like this! I haven’t checked if GPT has any circuits that fit this form; from what I’ve read of the Transformer Circuits sequence they don’t seem to have found predicted tokens in the residual streams. The activation space gradient descent theory is equally compelling, and equally unproven. Someone (you? me? anthropic?) should poke around in the weights of an LLM and see if they can find something that looks like this.

No convincing evidence for gradient descent in activation space

Blaine12 Apr 2023 4:48 UTC

86 points

9 comments20 min readLW link

Blaine

Tokyo AI Safety 2025: Call For Papers

We ran an AI safety con­fer­ence in Tokyo. It went re­ally well. Come next year!

Last call for sub­mis­sions for TAIS 2024!

An­nounc­ing TAIS 2024

No con­vinc­ing ev­i­dence for gra­di­ent de­scent in ac­ti­va­tion space

We ran an AI safety conference in Tokyo. It went really well. Come next year!

Last call for submissions for TAIS 2024!

Announcing TAIS 2024

No convincing evidence for gradient descent in activation space