RSS

Multi-Agent Safety

TagLast edit: 9 Feb 2026 12:45 UTC by Hiroshi Yamakawa

The In­her­i­tance Thresh­old: do be­hav­ioral norms sur­vive a hand­off be­tween LLM agents?

Viktor Trncik13 Jul 2026 17:32 UTC
1 point
0 comments3 min readLW link
(doi.org)

The Multi-Agent Minefield: Can LLMs Co­op­er­ate to Avoid Global Catas­tro­phe?

Isabel Dahlgren16 Feb 2026 18:53 UTC
1 point
0 comments5 min readLW link

Co­er­cion and De­cep­tion in AI-to-AI Management

10 Aug 2026 16:13 UTC
12 points
0 comments8 min readLW link
(compassionalignedml.substack.com)

I’m 18, Failed Chem­istry, and I Think I Found Some­thing in the Align­ment Problem

Vansh Ahuja16 May 2026 20:25 UTC
1 point
0 comments1 min readLW link

A Multi-Agent Ex­ten­sion for Petri

carissacullen22 Jul 2026 21:51 UTC
10 points
0 comments4 min readLW link

Han­ing Align­ment Pro­to­col: Emer­gent Hu­man-Com­pat­i­ble Values in Hy­brid Multi-Agent En­vi­ron­ments (A Con­cep­tual Pro­posal)

Josh Haning25 Jun 2026 20:45 UTC
1 point
0 comments1 min readLW link

The Multi-Agent Minefield: Can LLMs Co­op­er­ate to Avoid Global Catas­tro­phe?

17 Feb 2026 16:55 UTC
15 points
2 comments5 min readLW link
No comments.