TL;DR — Microsoft announced Project Perception on July 27 and it entered public preview on August 3, delivered inside Microsoft Defender. It coordinates red, blue and green AI agents on a new “cyber stack,” powered by a purpose-built model, MAI-Cyber-1-Flash. The genuine shift is from finding issues to acting on them — though autonomous remediation is deliberately held back from the initial preview. Below: what shipped, what’s gated, how the architecture fits together, and the questions worth testing before you lean on it.
On July 27, Microsoft announced Project Perception, an agentic cybersecurity platform built for a world where both attackers and defenders operate through autonomous AI. One week later, on August 3, it entered public preview inside Microsoft Defender. The premise is direct: security models built around human-scale triage can’t simply be accelerated enough to answer offense that now moves at machine speed. So rather than adding another assistant to the SOC, Microsoft is shipping a system meant to perceive risk, reason over it, and take protective action — while keeping a human in the loop for anything consequential.
That framing matters, because the real story is a shift in ambition. Much of what the market has shipped so far is good at finding — surfacing vulnerabilities. Finding is not defending. Project Perception’s bet is that the value lives in closing the loop from a finding to a fix without a hand-off at every step.
The three-color agent model
At the core of Perception is a workforce of specialized agents organized by role. Red agents probe like an attacker, mapping attack paths and exposures. Blue agents investigate what red surfaces and decide which findings represent meaningful risk. Green agents take corrective action: remediate, patch and harden. An orchestrator routes work between them and a message bus passes context, so a finding can flow into an investigation and then into a fix as a continuous loop rather than a stack of tickets.

Figure 1: Red probes, blue triages, green fixes. An orchestrator routes across a shared message bus; a human-approval gate sits in front of consequential green actions; the loop runs continuously.
Identity and governance for the agents themselves run through Microsoft Agent 365 and Entra Agent ID, and human approval is required for consequential actions. Microsoft says a single one of these workflows can compress roughly 146 hours of specialist time — a figure aimed squarely at teams dealing with alert fatigue.
A new cyber stack — and a purpose-built model
Microsoft frames Perception as more than agents bolted onto existing products. It describes a purpose-built cyber stack running end to end: sensors collect telemetry; a security-context layer distills that flood into a token-efficient graph agents can actually reason over; models supply intelligence; a harness coordinates models and agents; agents apply that intelligence; and actuators translate decisions into enforced protection. The context graph is the quiet centerpiece — no system reasons directly over the trillions of signals a large enterprise generates daily, so distilling them into a navigable graph is what makes the whole thing tractable.

Figure 2: The six layers of the cyber stack, from signals and sensors down to actuators. The red/blue/green agents sit at layer 05.
Underneath sits MAI-Cyber-1-Flash, Microsoft’s first in-house model built specifically for cybersecurity, initially focused on software vulnerability analysis. It’s designed to handle the bulk of the workload cheaply and route only the hardest slice — Microsoft pegs it at roughly 10% — to a frontier model (OpenAI’s GPT-5.4) reserved for exceptionally difficult tasks. Microsoft reports the configuration scored 96% on the CyberGym benchmark at roughly half the cost of its previous MDASH setup. This multi-model routing — right model, right problem, right cost — is the operating philosophy of the platform.

Figure 3: Microsoft’s stated routing design. The 90/10 split and the 96% CyberGym figure are Microsoft’s own numbers, not independently verified.
What’s actually in the preview — and what isn’t
This is the part worth reading carefully before you brief anyone on it. Autonomous remediation is not in the initial preview. Microsoft is staging it: reversible actions like isolating a device are slated for later this year, while riskier moves like patching a production host stay behind human approval. The green agents can propose and, over time, execute fixes — but the mandate to let them act autonomously is being handed out deliberately.
That gating is a sensible call, and it isn’t unique — comparable systems that rewrite vulnerable code still hand a validated patch to a human before deployment. The hard problem in agentic security was never generating a plausible fix. It’s earning enough trust to let an agent execute one against your environment.
Separately, Microsoft also expanded runtime protection for AI agents governed through Agent 365 — real-time threat detection for agents built with Microsoft tooling as well as supported third-party and locally running agents, with local runtime protection on Windows endpoints available in preview. It’s a reminder that Microsoft is playing both sides of the agentic wave: using agents to defend, and defending the agents enterprises are now deploying everywhere.
Pricing: consumption, with an asterisk
Perception is offered as a consumption-based service measured in Security Compute Units (SCUs), with different agents consuming SCUs at different rates. Microsoft has not published numeric SCU rates. For anyone modeling cost, that’s the open question: agentic loops run on tokens, and a system that reasons continuously over a large estate can add up quickly. Whether bills stay predictable at production scale is the adoption question the cost-efficiency story is meant to answer.
Agents are table stakes; presence is the moat
Here’s the part worth being clear-eyed about. Agents aren’t really Microsoft’s differentiator. Every large platform is shipping some version of agentic security, and that capability is fast becoming table stakes.
Microsoft’s real edge isn’t the agents. It’s that Defender already sits where its customers live — Entra identity, Windows endpoints, and the management plane most organizations already run through.
Fold agentic discovery and response into Defender, price it by consumption, and you deliver much of what specialist tools offer — from a first-party seat inside the estate customers already operate. That’s the classic platform move: single-function tooling that lives entirely inside the Microsoft world tends to become a feature rather than a category. It isn’t automatic displacement — depth and independent proof still matter, and reach isn’t the same thing as efficacy — but the gravity is real.
Questions worth testing in your own environment
- Benchmarks aren’t outcomes. A 96% CyberGym score demonstrates capability in a test harness, not defensive results in your environment. There’s still no cheap, independent way to verify real-world efficacy — and the headline numbers are Microsoft’s own.
- The graph is only as good as its coverage. The context graph is the moat, but an asset-and-identity model is only as trustworthy as what it can see. Shadow IT, unmanaged identities and stale inventory — exactly where real intrusions start — are the blind spots that decide outcomes, not model quality.
- The demo is code-shaped; attacks often aren’t. The flagship walkthrough runs an attack path through a web app — SQL injection, cross-site scripting, a tidy remediation plan, all expressible in code. Many real intrusions aren’t: credential harvesting and help-desk social engineering don’t reduce to a code fix. Worth watching how far red and blue extend into identity abuse.
- Autonomy concentrates blast radius. Delegating to agents that chain steps and reach production concentrates trust. Governing agents that act is harder than governing a single model — something to weigh against your own risk tolerance as the roadmap adds autonomy.
If you’re on the defending side
Perception is worth piloting for the triage-to-remediation compression alone — but treat the preview as exactly that. What you can rely on today is faster investigation and prioritized findings; the autonomous-fix promise is a roadmap item, not a shipped capability. Three practical things to do during the preview window: validate how much of your real estate the context graph actually covers, especially non-Microsoft and unmanaged assets; model SCU consumption against your own alert volume before it becomes a line item; and decide your own gating policy for green-agent actions rather than inheriting the default. The platform will keep pace with the market on features. Whether it earns the trust to act in your environment is a decision you still own.
Sources & notes: Microsoft’s official announcement and Project Perception product page, plus reporting and analysis from GeekWire, Redmond Magazine, Windows Forum, The Elec and Futurum Group. Benchmark figures and the 146-hour, ~50% cost and 90/10 routing claims are Microsoft’s own and remain independently unverified at the time of writing. Diagrams are original, created for this post. This is an independent analysis and not affiliated with or endorsed by Microsoft.