John Hughes

Karma: 569

Former MATS scholar working on scalable oversight and adversarial robustness.

Why Do Some Language Models Fake Alignment While Others Don’t?

8 Jul 2025 21:49 UTC

152 points

(arxiv.org)

8 Apr 2025 17:32 UTC

146 points

John Hughes 22 Jan 2025 10:34 UTC
3 points
0
in reply to: Jan Betley’s comment on: Tips and Code for Empirical Research Workflows
Threads are managed by the OS and each thread has an overhead in starting up/switching. The asyncio coroutines are more lightweight since they are managed within the Python runtime (rather than OS) and share the memory within the main thread. This allows you to use tens of thousands of async coroutines, which isn’t possible with threads AFAIK. So I recommend asyncio for LLM API calls since often, in my experience, I need to scale up to thousands of concurrents. In my opinion, learning about asyncio is a very high ROI for empirical research.

20 Jan 2025 22:31 UTC

95 points

8 Jan 2025 5:06 UTC

93 points

John Hughes 18 Dec 2024 11:07 UTC
3 points
0
in reply to: anaguma’s comment on: Best-of-N Jailbreaking
For o1-mini, the ASR at 3000 samples is 69.2% and has a similar trajectory to Claude 3.5 Sonnet. Upon quick manual inspection, the false positive rate is very small. So it seems the reasoning post training for o1-mini helps with robustness a bit compared to gpt-4o-mini (which is nearer 90% at 3000 steps). But it is still significantly compromised when using BoN...

John Hughes 14 Dec 2024 23:58 UTC
3 points
0
in reply to: anaguma’s comment on: Best-of-N Jailbreaking
Thanks! Running o1 now, will report back.

John Hughes 14 Dec 2024 23:47 UTC
6 points
0
in reply to: Thomas Kwa’s comment on: Best-of-N Jailbreaking
Circuit breaking changes the exponent a bit if you compare it to the slope of Llama3 (Figure 3 of the paper). So, there is some evidence that circuit breaking does more than just shifting the intercept when exposed to best-of-N attacks. This contrasts with the adversarial training on attacks in the MSJ paper, which doesn’t change the slope (we might see something similar if we do adversarial training with BoN attacks). Also, I expect that using input-output classifiers will change the slope significantly. Understanding how these slopes change with different defenses is the work we plan to do next!