RSS

wesg

Karma: 878

Working on interpretability.

Find out more here: https://​​wesg.me/​​

A global workspace in lan­guage models

wesg6 Jul 2026 18:04 UTC
366 points
66 comments17 min readLW link
(www.anthropic.com)

Re­fusal in LLMs is me­di­ated by a sin­gle direction

27 Apr 2024 11:13 UTC
260 points
96 comments10 min readLW link1 review

SAE re­con­struc­tion er­rors are (em­piri­cally) pathological

wesg29 Mar 2024 16:37 UTC
108 points
16 comments8 min readLW link

Find­ing Neu­rons in a Haystack: Case Stud­ies with Sparse Probing

3 May 2023 13:30 UTC
33 points
6 comments2 min readLW link1 review
(arxiv.org)