3 lessons from frontier AI vulnerability research
Microsoft's FORGE Lab uses AI to find vulnerabilities in Windows and Linux.
Summary
Microsoft Security's FORGE Lab is leveraging AI for autonomous vulnerability research, focusing on scaling discovery and remediation. The lab has identified 140 CVEs in Windows and contributed to open-source projects like the Linux kernel. Key lessons learned emphasize the shift from frontier capability to scalable processes, efficient reasoning economics, and coordinated validation and remediation.
Full text
Share Link copied to clipboard! TagsVulnerabilityWindowsContent typesResearchTopicsAI and agents The mission of Microsoft Security’s Frontier Offensive Research & Generative Exploitation (FORGE) Lab is to advance the frontier of autonomous security engineering. We’re building a team that enables AI-native vulnerability research at Microsoft, pushing the boundaries of finding and fixing zero-day vulnerabilities. Four principles guide that work: autonomy over labor, defense through offense, building ecosystems over individual examples, and understanding over findings. This post reflects those principles in practice across Windows, the Linux kernel, and widely used open-source projects. Learn more about the Microsoft FORGE Lab From May 2026 through September 2026, FORGE helped discover Windows vulnerabilities assigned 140 common vulnerabilities and exposures (CVEs), including 52 addressed in September 2026’s security release alone. The work reaches well beyond Windows. FORGE members submitted 155 internally validated reports across 23 open-source projects, including the Linux kernel. And through Akrites, a Linux Foundation initiative that coordinates confidential remediation and disclosure for vulnerabilities in critical open-source software, one of our Linux reports became the first Akrites submission to result in a patch merged into the Linux kernel. These results demonstrate that agentic discovery can operate at meaningful volume, but they also expose the next constraint. Finding a difficult vulnerability proves an agent’s frontier capability. Repeatedly converting such findings into security updates requires a different kind of system—one that can move credible work through validation, remediation, and release. Discovery creates security value only when validation and remediation can keep pace. For security leaders, the question is no longer only whether AI can find vulnerabilities, but whether an organization can validate and fix them as fast as they are found. Three lessons follow, each marking a shift in how we approach vulnerability research: From frontier capability to scale. From token consumption to reasoning economics. From isolated discovery to coordinated validation and remediation. The scarce resource is not always model intelligence. It can be a working build and deployment, a reproducible trigger, or even an engineer’s time. At FORGE Lab, we use the multi-model agentic scanning harness, codename MDASH, to organize this work. Our May 2026 introduction detailed its design, while the June 2026 update covered pipeline improvements and benchmark analysis. Here, we focus on what those experiences teach us about operating vulnerability research at scale—across Windows and open-source projects. 140 CVEs addressed through Microsoft Patch Tuesday since May 2026. Monthly totals reflect announcement and servicing cohorts, not discovery dates or scan throughput. Source: Microsoft FORGE Lab study, May 2026 to September 2026. Shift 1. From frontier capability to scale The frontier question is whether a system can find a bug that demands deep reasoning, as we explored in our early work on the CyberGym benchmark. The scale question is whether it can repeat that result across targets without rebuilding every environment, investigation, and review process from scratch. Better models still matter, especially for difficult or unfamiliar bug classes. But once discovery becomes repeatable, model capability is only one constraint on useful output. Adding auditors can increase candidate volume without increasing the rate of validated findings or shipped fixes. If reports arrive faster than reviewers—in our case, Microsoft Security Response Center (MSRC)—can resolve them, the immediate result is a growing queue. A larger search budget can even reduce useful throughput when duplicate or poorly supported reports consume attention that stronger findings need. The unit to optimize, therefore, is neither the scan nor the report. It is a reproducible finding that advances with enough evidence for the next stage to act. Reusable preparation, deduplication, and review capacity are therefore core components of the research system. As agentic systems produce more candidates, the bottleneck shifts from discovery to determining which reports are real, reachable, and security-relevant. Project-specific automated provers—such as proof of vulnerability (PoV)/proof of concept (PoC) generators, harness builders, and trigger-input finders—turn plausible reports into reproducible evidence that engineers and maintainers can act on. By finding crashing inputs, confirming reachable execution paths, and producing regression-ready triggers within each project’s build and test environment, these tools can help reduce human triage, accelerate remediation, and make vulnerability research sustainable at scale. For example, by leveraging deterministic algorithms such as abstract syntax tree (AST), one of our internal projects reduced about 45% duplicate findings across multiple scans of the same code which reduced the load on the PoC generator and subsequently human triage effort. More candidates do not create more throughput when the review queue grows faster than fixes ship. Shift 2. From token consumption to reasoning economics At scale, every repeated orientation, speculative branch, and redundant debate carries a cost. But minimizing tokens alone is the wrong objective. A short, ambiguous report may be cheap to generate yet expensive to investigate; a longer analysis that establishes the missing execution path may reduce total system cost. The key question is where the next unit of reasoning will change a decision. If an index can identify callers, asking a frontier model to rediscover them wastes capacity. If the uncertainty is whether two lifetime conditions can coexist across callbacks, deeper reasoning may be warranted. If the question is whether an input triggers the failure, an executable check can provide evidence that more prose cannot. MDASH combines frontier and distilled models, specialized auditors, and code-analysis tools, enabling work to be allocated by task. That flexibility does not guarantee optimal allocation. Routing routine work to cheaper models, reusing verified context, and escalating unresolved questions to stronger reasoning are hypotheses to test—not efficiency gains to assume. A useful allocation policy starts by identifying what remains unknown: a caller, a build configuration, a reproducer, or a causal explanation. The next action should close that specific evidence gap. Repeating a review without adding evidence spends more tokens while preserving the same uncertainty. Early filtering can also discard real bugs, so any savings must be measured against coverage and missed findings. Scan outcomes can also become training data. At scale, vulnerability scanning should become a training loop, not only a discovery pipeline. Each MDASH run produces signals that can improve future models: true-positive and false-positive verdicts, duplicate findings, failed reachability claims, reviewer feedback, verifier results, severity assessments, patch outcomes, and regression-test evidence. Capturing those signals with the code context and causal argument behind each candidate creates the dataset needed for reinforcement learning and fine-tuning specialized cyber models. The goal is for future models to learn not just what a vulnerability looks like, but which findings survive validation, which explanations help humans and provers act, and which patterns lead to useful remediation. Spend reasoning to remove uncertainty, not simply to produce more analysis. Shift 3. Validation and remediation as a continuous learning loop Validation and remediation are not downstream cleanup; they are part of the discovery loop. Each candidate should move through automated verification, human review, patch development, and regression testing, with every stage returning evidence to the sys