“Why Not Just...”johnswentworth8 Aug 2022 18:15 UTCA compendium of rants about alignment proposals, of varying charitability.Deep Learning Systems Are Not Less Interpretable Than Logic/Probability/Etcjohnswentworth4 Jun 2022 5:41 UTC170 points57 comments2 min readLW link1 reviewGodzilla Strategiesjohnswentworth11 Jun 2022 15:44 UTC176 points72 comments3 min readLW linkRant on Problem Factorization for Alignmentjohnswentworth5 Aug 2022 19:23 UTC112 points53 comments7 min readLW linkInterpretability/Tool-ness/Alignment/Corrigibility are not Composablejohnswentworth8 Aug 2022 18:05 UTC151 points13 comments3 min readLW linkHow To Go From Interpretability To Alignment: Just Retarget The Searchjohnswentworth10 Aug 2022 16:08 UTC215 points34 comments3 min readLW link1 reviewOversight Misses 100% of Thoughts The AI Does Not Thinkjohnswentworth12 Aug 2022 16:30 UTC126 points49 comments1 min readLW linkHuman Mimicry Mainly Works When We’re Already Closejohnswentworth17 Aug 2022 18:41 UTC83 points16 comments5 min readLW linkWorlds Where Iterative Design Failsjohnswentworth30 Aug 2022 20:48 UTC244 points32 comments10 min readLW link1 reviewWhy Not Just… Build Weak AI Tools For AI Alignment Research?johnswentworth5 Mar 2023 0:12 UTC188 points18 comments6 min readLW linkWhy Not Just Outsource Alignment Research To An AI?johnswentworth9 Mar 2023 21:49 UTC162 points50 comments9 min readLW link1 reviewOpenAI Launches Superalignment TaskforceZvi11 Jul 2023 13:00 UTC150 points40 comments49 min readLW link(thezvi.wordpress.com)Why Not Just Train For Interpretability?johnswentworth21 Nov 2025 22:08 UTC59 points12 comments4 min readLW link