Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
fbarez
Karma:
132
All
Posts
Comments
New
Top
Old
Martian Interpretability Challenge: The Core Problems In Interpretability
fbarez
11 Mar 2026 17:41 UTC
9
points
0
comments
9
min read
LW
link
Automated Interpretability-Driven Model Auditing and Control: A Research Agenda
fbarez
12 Jan 2026 19:55 UTC
9
points
0
comments
1
min read
LW
link
Best-of-N Jailbreaking
John Hughes
,
saraprice
,
Aengus Lynch
,
Rylan Schaeffer
,
fbarez
,
Henry Sleight
,
Ethan Perez
and
mrinank_sharma
14 Dec 2024 4:58 UTC
79
points
5
comments
2
min read
LW
link
(arxiv.org)
Visualizing neural network planning
Nevan Wichers
,
Victor Tao
,
fbarez
and
Riccardo Volpato
9 May 2024 6:40 UTC
4
points
0
comments
5
min read
LW
link
Mechanistic Interpretability Workshop Happening at ICML 2024!
Neel Nanda
,
LawrenceC
and
fbarez
3 May 2024 1:18 UTC
48
points
6
comments
1
min read
LW
link
Automated Sandwiching & Quantifying Human-LLM Cooperation: ScaleOversight hackathon results
Esben Kran
,
fbarez
,
Sabrina Zaki
,
gabrielrecc
and
rz2383
23 Feb 2023 10:48 UTC
8
points
0
comments
6
min read
LW
link
Back to top