topic
Security Research
2 posts tagged “Security Research”.
Once a prompt injection reaches the AI, does it obey? I tested that too
Part 2 of a defensive-research series: once a planted instruction reaches an AI SOC-triage assistant, how often does it obey? Measured across two models and four defence conditions — with the finding that instructing the reader beats fancy ingestion tags, and that tags can backfire on cheap models.
Do prompt injections survive the Microsoft Sentinel pipeline? I measured it.
A defensive-research experiment: if an attacker hides an instruction inside a log field, does it survive a real Microsoft Sentinel ingestion pipeline all the way to where an AI assistant would read it? The answer — and why the connector's transform is the real control surface.