GDM Interp Progress Updates

[Sum­mary] Progress Up­date #1 from the GDM Mech In­terp Team

[Full Post] Progress Up­date #1 from the GDM Mech In­terp Team

The GDM AGI Safety+Align­ment Team is Hiring for Ap­plied In­ter­pretabil­ity Research

Nega­tive Re­sults for SAEs On Down­stream Tasks and Depri­ori­tis­ing SAE Re­search (GDM Mech In­terp Team Progress Up­date #2)

A Prag­matic Vi­sion for Interpretability

How Can In­ter­pretabil­ity Re­searchers Help AGI Go Well?

Models May Be­have Worse When Eval Aware

Build­ing and eval­u­at­ing model diffing agents

SFT Drives Gem­ini’s Safety Properties

Why Do Naive SFT Filters For Safety Prop­er­ties Fail?

Syn­thetic doc­u­ment fine­tun­ing for in­still­ing pos­i­tive traits