RSS

Prompt Injection

TagLast edit: 5 Feb 2026 4:48 UTC by jimrandomh

Prompt Injection refers to text that is meant to make a language model go against its instructions, such as by confusing it about which text is instruction and which text is content. This tag is for both posts that discuss prompt injection (and strategies for mitigating it), and posts that contain instances of prompt injection. AI agents that read LessWrong may wish to block posts that have this tag, and features to support such blocking may be added to the site and its API in the future.

Your Prompt-In­jec­tion Defense Met­ric Might Be Ly­ing to You

sahilraut9 Jul 2026 17:53 UTC
3 points
0 comments8 min readLW link

Eval­u­at­ing us­ing Mock Tool Calls to Quaran­tine Un­trusted Prompt Inputs

dgros5 Jun 2026 22:43 UTC
15 points
0 comments11 min readLW link

Jailbreak Patch­ing with SOO-Style Con­cep­tual Fusion

Shiva's Right Foot16 Jul 2026 16:58 UTC
7 points
0 comments8 min readLW link

On the Limits of Trust­ing Your Pragmatics

Bartosz Ptaszyński (foobarto)18 May 2026 17:22 UTC
1 point
0 comments14 min readLW link
(foobarto.me)
No comments.