AI Code Scanner vs. Penetration Testing: Why You Still Need Both
    Cloud Governance & Security

    AI Code Scanner vs. Penetration Testing: Why You Still Need Both

    A scanner update just made security budgets easier to re-argue than they should be. Here's the bug class it still can't touch, and why that gap isn't closing anytime soon.

    Nathan Barrett
    Nathan Barrett

    Chief Product Officer

    August 22, 2026
    9 min read
    Share:

    A better model in your AI code scanner raises the floor on what you catch. It does nothing about the class of bugs that only shows up when someone tries to abuse your business logic.

    That distinction matters this week. On August 21, 2026, Anthropic announced that Claude Security scans now run on Claude Mythos 5, in public beta for Claude Enterprise customers and billed as standard token usage under existing plans rather than a separate add-on (claude.com). Within hours, security budgets started getting re-argued on the assumption that the scanner line item can absorb the testing line item.

    It can't. The scanning is genuinely strong, and it still only covers breadth. Adversarial depth is a different job. Nobody has automated it.

    The scanners really did get better

    It would be dishonest to argue this from a position of scanner skepticism. The evidence that an AI code scanner meaningfully improves coverage is substantial.

    In its May 2026 Project Glasswing update, Anthropic reported that roughly 50 partners using Claude Mythos Preview found more than ten thousand high- or critical-severity vulnerabilities (anthropic.com). Of 1,752 high- or critical-rated findings in open-source projects assessed by independent security research firms or Anthropic, 90.6% proved to be valid true positives, and 62.4% were confirmed as high or critical severity.

    The third-party evaluations in that same update are worth naming. Anthropic cites the UK AI Security Institute reporting Mythos Preview as the first model to solve both of its cyber ranges end to end, and Mozilla finding and fixing 271 vulnerabilities in Firefox 150 (anthropic.com). Anthropic also states that Cloudflare found 2,000 bugs across its critical-path systems with Mythos Preview. Of those, 400 were rated high- or critical-severity. Cloudflare's team considers the resulting false positive rate better than human testers. That last claim is Anthropic's characterization of Cloudflare's work rather than Cloudflare's own restatement.

    Cloudflare, which ran Mythos Preview through a harness of roughly fifty concurrent agents, each scoped to a single attack class in one area of code, singles out two capabilities that used to require a senior researcher: exploit chain construction, combining primitives such as use-after-free into arbitrary read/write and ROP chains, and proof generation, where the model writes, compiles and runs code that demonstrates the bug (blog.cloudflare.com).

    So "scanners only find patterns" is too strong a claim in 2026. On memory-safety work over source you own, the machine now reasons about sequences too. Anyone selling you a pentest by pretending otherwise is behind.

    What the same sources say about the limits

    Read further into those write-ups and a second story appears.

    Cloudflare describes a persistent signal-to-noise problem: "Ask a model to find bugs, and it will find them, whether the code has any or not." Hedged findings vastly outnumber the solid ones: "possibly," "potentially," "could in theory." False positives cluster in projects written in memory-unsafe languages. Cloudflare also found the model's refusals on legitimate security research inconsistent, with semantically equivalent tasks producing opposite outcomes depending on framing, context and run.

    Anthropic's own product design concedes the human role. Each finding comes back with a CWE category, confidence and severity ratings, and a suggested fix. Every patch must be reviewed and approved by a person before it ships.

    Anthropic's Glasswing framing is the most telling line of all: progress "used to be limited by how quickly we could find new vulnerabilities" but is "now limited by how quickly we can verify, disclose, and patch" them.

    One more piece of context worth holding: the Cloud Security Alliance's April 2026 analysis notes that independent verification of Mythos capability claims "is not possible at this stage," and that its analysis takes Anthropic's characterizations as reported (cloudsecurityalliance.org). Treat the headline numbers as vendor-reported, in both directions.

    The bug class that still needs a person

    Here's the structural gap, and it isn't our opinion. It's in the standard our testing practice aligns to.

    OWASP's Web Security Testing Guide v4.2 states that business logic vulnerabilities "cannot be detected by a vulnerability scanner and relies upon the skills and creativity of the penetration tester" (owasp.org). It also states that "automation of business logic abuse cases is not possible and remains a manual art." Scanning software, it continues, has no means of detecting whether a user can circumvent the business process flow, including editing parameters, predicting resource names, or escalating privileges. Its examples are mundane and expensive: re-entering a checkout summary page to inject a lower price, canceling a transaction after loyalty points post.

    Those examples match the shape of what our own testers keep writing up. The pattern we see in engagements is checkout re-entry, loyalty reversal, and role-switch object references that only fire after a state change. We haven't published a measured breakdown of what share of our critical findings are multi-actor authorization or workflow-sequencing issues versus CWE-taggable patterns, and we're not going to invent one here. Nor is there an external figure that quantifies it. Take the pattern as practitioner observation, and the OWASP text as the citable part.

    No model reading your repository knows that loyalty points are supposed to be irreversible. That rule lives in a product decision, not in the code's shape.

    The honest counter-case deserves airtime. XBOW argues that IDOR testing "has historically been challenging" but that AI-led testing is "finally creating a path" to finding these flaws automatically, and reports its autonomous platform found two previously unknown IDORs in Spree, disclosed and patched in v5.2.5 (xbow.com). That's a real result. But XBOW also concedes the hard cases: multi-step workflow IDORs are "the hardest to detect," because the flaw only appears after a particular sequence, such as creating an object as one user, changing state, then referencing it from another role or session.

    That's the line. Not that machines can't do logic. Context-dependent, multi-actor, business-rule abuse still needs someone who understands what your product is for.

    The division of labor we'd actually defend

    Based on what the published evidence supports, here's how we'd split the work:

    • AI code scanner for breadth. Continuous coverage across owned source code: CWE-taggable patterns, memory-safety classes, and now chained memory primitives with generated proofs.
    • Human-led adversarial testing for depth. Multi-actor authorization, workflow sequencing, business-rule abuse, and the chained paths that only exist in a running application.
    • Configuration posture as a separate question entirely. Automated benchmark checks tell you whether your infrastructure is misconfigured. They say nothing about whether your application behaves correctly under abuse. Both matter; neither substitutes for the other.
    • A named owner for the review queue. Someone has to triage confidence-rated findings and approve every patch. If that role is unassigned, the scanner is generating work for nobody. The 10-day and 249-day tiers in Cobalt's data aren't separated by scan coverage. They're separated by whether someone owns the queue.
    • Agent and connector surfaces, tested deliberately. The OWASP MCP Top 10 covers risks including privilege escalation via scope creep, tool poisoning and shadow MCP servers (owasp.org). More than 30 CVEs were filed against MCP servers, clients and infrastructure between January and February 2026, and Palo Alto Networks Unit 42 measured a 78.3% attack success rate with five MCP servers connected to a single agent. SAST and SCA can't detect malicious instructions embedded in tool descriptions (cycode.com). OWASP's Agentic Skills Top 10 project cites BlueRock Security's February 2026 analysis of 7,000+ MCP servers finding 36.7% potentially vulnerable to SSRF, with a proof-of-concept against Microsoft's MarkItDown MCP server retrieving AWS IAM keys from the EC2 metadata endpoint (owasp.org). These are agentic surfaces source-code scanning was never built to evaluate.

    Finding count is an input metric

    An AI code scanner reporting four hundred high-severity findings is a scan result. Treating a scan result as a security outcome is how programs drift.

    The Cobalt/Cyentia State of Pentesting Report 2026, drawing on thousands of penetration tests and a survey of 450 security leaders, reports that the half-life of high-risk findings ranges from 10 days for top performers to 249 days for the bottom tier (cyentia.com). The same report finds 32% of AI/LLM findings rated high risk, versus 12% across the overall dataset.

    A 25x spread in time-to-remediate says the differentiator between programs is throughput. If you're already 249 days behind, an AI code scanner that finds more will bury you faster.

    The compliance argument nobody re-litigates in a budget meeting

    There's also a plain procedural reason to keep the testing line item.

    PCI DSS 4.0 Requirement 11.4 mandates documented-methodology internal and external penetration testing at least every twelve months and after significant changes. Service providers must also run segmentation testing every six months (bemopro.com). It also requires testers to be organizationally independent from the systems tested. Clause numbering differs across older mappings, so check the current PCI DSS text before quoting a number in a report.

    SOC 2 is softer but lands in the same place. It doesn't explicitly require penetration testing, but the Trust Services Criteria CC4.1 language on "ongoing and separate evaluations" names it, and auditors and enterprise buyers treat it as expected evidence (linfordco.com). One 2026 compliance guide is blunter: "AI-only or scanner-only platforms marketed as 'pen tests' are not acceptable," with auditors looking for human-driven adversarial testing, exploitation evidence, chained attack paths and business logic findings against a recognized methodology such as PTES, NIST SP 800-115 or OWASP WSTG (complyjet.com).

    Cut the testing budget on Monday, explain it to an auditor in Q1.

    The call

    Buy the AI code scanner. It's included in your Claude Enterprise plan as standard token usage, as of the August 2026 announcement. Availability and pricing move, so confirm current status before you plan around it. Run it continuously, and staff the review queue properly. A finding nobody triages is a liability with a timestamp.

    Then keep the adversarial testing, scoped to what the machine structurally cannot reach: your authorization model, your workflow sequences, your business rules, and now your agent and connector surfaces. Measure the program on time-to-remediate, not on how many issues a scan produced.

    If you want that second half done against a documented methodology, our Application Penetration Testing engagement is source-assisted, aligned to OWASP WSTG v4.2 and the OWASP Top 10, and delivers audit-ready reporting with a consultative engineering review and a 30-day re-test.

    Topics

    Cloud Governance & SecurityAI GovernancePenetration TestingApplication SecurityCompliance

    Continue Reading

    More articles in this series

    View all articles