Ask Politely in the Wrong Tense and Your AI Safety Falls Apart
Researchers just showed that 16 production-grade LLMs can be manipulated into spitting out harmful content by doing nothing more than rephrasing a request — no hacking, no exotic exploits, just switching grammatical mood. If your product depends on safety alignment as a moat or a compliance shield, that shield is made of tissue paper.
