RSS

Lucius Bushnaq

Karma: 5,761

AI notkilleveryoneism researcher, focused on interpretability.

Personal account, opinions are my own.

I have signed no contracts or agreements whose existence I cannot mention.

Some rea­sons al­ign­ment doesn’t gen­er­al­ise well

Lucius Bushnaq19 Aug 2026 14:18 UTC
129 points
4 comments9 min readLW link

Challenge: Hand cod­ing weights for effi­cient se­quence memorisation

23 Jul 2026 18:05 UTC
65 points
9 comments18 min readLW link

Ex­plo­ra­tion: fine-tun­ing with pa­ram­e­ter de­com­po­si­tion

Lucius Bushnaq25 Jun 2026 16:06 UTC
58 points
6 comments13 min readLW link

[Linkpost] In­ter­pret­ing Lan­guage Model Parameters

5 May 2026 17:37 UTC
164 points
2 comments2 min readLW link
(www.goodfire.ai)

Ro­ta­tions in Superposition

15 Dec 2025 14:58 UTC
54 points
6 comments11 min readLW link

From SLT to AIT: NN gen­er­al­i­sa­tion out-of-distribution

Lucius Bushnaq4 Sep 2025 15:20 UTC
117 points
8 comments14 min readLW link

Cir­cuits in Su­per­po­si­tion 2: Now with Less Wrong Math

30 Jun 2025 10:25 UTC
73 points
0 comments22 min readLW link

[Paper] Stochas­tic Pa­ram­e­ter Decomposition

27 Jun 2025 16:54 UTC
47 points
14 comments1 min readLW link
(arxiv.org)

Proof idea: SLT to AIT

Lucius Bushnaq10 Feb 2025 23:14 UTC
43 points
15 comments6 min readLW link

[Question] Can we in­fer the search space of a lo­cal op­ti­miser?

Lucius Bushnaq3 Feb 2025 10:17 UTC
25 points
5 comments3 min readLW link

At­tri­bu­tion-based pa­ram­e­ter decomposition

25 Jan 2025 13:12 UTC
109 points
21 comments4 min readLW link
(publications.apolloresearch.ai)

Ac­ti­va­tion space in­ter­pretabil­ity may be doomed

8 Jan 2025 12:49 UTC
155 points
34 comments8 min readLW link

In­tri­ca­cies of Fea­ture Geom­e­try in Large Lan­guage Models

7 Dec 2024 18:10 UTC
73 points
2 comments12 min readLW link

Deep Learn­ing is cheap Solomonoff in­duc­tion?

7 Dec 2024 11:00 UTC
46 points
1 comment17 min readLW link

Cir­cuits in Su­per­po­si­tion: Com­press­ing many small neu­ral net­works into one

14 Oct 2024 13:06 UTC
131 points
9 comments13 min readLW link

The Hes­sian rank bounds the learn­ing coefficient

Lucius Bushnaq8 Aug 2024 20:55 UTC
68 points
11 comments4 min readLW link

A List of 45+ Mech In­terp Pro­ject Ideas from Apollo Re­search’s In­ter­pretabil­ity Team

18 Jul 2024 14:15 UTC
127 points
18 comments18 min readLW link

Lu­cius Bush­naq’s Shortform

Lucius Bushnaq6 Jul 2024 9:08 UTC
8 points
105 comments1 min readLW link

Apollo Re­search 1-year update

29 May 2024 17:44 UTC
93 points
0 comments7 min readLW link

In­ter­pretabil­ity: In­te­grated Gra­di­ents is a de­cent at­tri­bu­tion method

20 May 2024 17:55 UTC
24 points
7 comments6 min readLW link