AI Governance Gaps in SOC 2 and ISO 27001 Audits
Auditing frameworks lack controls designed for AI systems that operate autonomously across vendors.

SOC 2 and ISO 27001 were finalized before agentic AI existed as a deployment pattern, and that timing explains almost everything about where the gaps sit today. SOC 2's Trust Services Criteria took their current form in 2017 [2][3][1][4]. ISO 27001 went through its last major refresh more recently still. Agentic AI, as something enterprises actually run in production, arrived in the eighteen months after that refresh. The standard itself never had the chance to see the thing it now gets asked to govern.
This is not a story about regulators falling behind out of neglect. If standards bodies need years to negotiate and reach consensus, governance frameworks will move slowly. That pace made sense for the technology it was built to cover. Agentic AI did not wait for consensus. It arrived fast, inside real companies, running real workflows, and the frameworks simply had no mechanism to absorb that speed.
Look at what changed on the ground. Enterprises now run large language model pipelines, and these cross multiple vendor boundaries in a single transaction. Retrieval-augmented generation systems reach out to third-party inference endpoints mid-request. Multi-agent workflows chain autonomous decisions together, and they do it all within the time it takes to answer one user prompt. The behavior that comes out the other end is emergent. It depends on context. It's distributed across systems that no single static configuration audit was ever built to see end to end.
A framework paper on AI Trust OS, written by researchers spanning Old Dominion University, Deloitte, Florida International University, and several industry labs, states the underlying problem directly: compliance methodologies built for deterministic, stateless web applications give auditors no way to discover, classify, or continuously check AI systems that grow organically across engineering teams without anyone formally signing off on them. That's the root of the mismatch. Everything the rest of this piece covers, from shadow AI tools to unmonitored vendor models to the improvisation auditors have resorted to, traces back to that one structural fact: the rulebook was finished before the thing it's supposed to rule existed.
What the Frameworks Test
SOC 2 is built around five Trust Services Criteria: Security, Availability, Processing Integrity, Confidentiality, and Privacy. Only Security is mandatory. The other four get scoped in or out at the organization's own discretion. That discretion creates an opening large enough to drive a fleet of AI products through: a company can leave Privacy out of scope entirely, even while its AI system is actively processing customer data, and still walk away with a clean report.
Four specific blind spots appear repeatedly in audit engagements, each traceable to a different piece of how these criteria were built.
Shadow AI assets are the first. Employees pick up AI tools through consumer interfaces, browser extensions, or quick API integrations, and IT often has no idea. That absence is invisible to a standard asset register. Since access-control and vendor-management criteria depend on that register existing in the first place, these tools are invisible to the audit by construction, not by oversight.
Unmonitored vendor models are the second. When an AI product sits on top of a third-party foundation model, somebody needs to check that vendor's own compliance status and how it retains data. No criterion in SOC 2 or ISO 27001 requires an auditor to do that check, so in practice, many engagements skip it.
Behavioral drift is the third. A model gets retrained, fine-tuned, or an upstream vendor pushes a silent update, so its outputs shift over time. Neither framework includes a control that calls for drift-monitoring dashboards or change-management tickets tied to model deployments, so a model can behave meaningfully differently in March than it did in January and nothing in the audit trail will register that change.
The fourth is the hardest one for a traditional auditor to work around: the agentic accountability gap. When an autonomous agent reads a file, writes data, calls an external tool, or triggers a connector without a discrete human request sitting in front of each action, something happened, but no identity-and-access control captured who authorized it. From an auditor's standpoint, that agent is a nonhuman identity. It initiates access events. It holds credentials. It can carry long-lived permissions that got provisioned quickly and never got reviewed again, and that runs directly against the least-privilege principle that SOC 2's CC6 criteria, especially CC6.1 and CC6.3, assume as a baseline. There's a subtler version of this same failure, too: an attacker with nothing more than read-only access can submit documents into a training dataset and introduce semantic manipulations that change system behavior later, without ever tripping a write-permission or data-integrity control. The check passes. The system is already compromised.
Researchers at UNSW built a coverage matrix, and it tested five major governance frameworks, ISO 27001 among them, against the specific failure causes AI systems produce. None of the five addressed shadow AI, speed asymmetry, or the governance vacuum at the level of operational detail the problem needs.
How auditors are improvising without AICPA guidance
The AICPA has not published AI-specific Trust Services Criteria. It has not issued AI-specific points of focus. No authoritative AI addendum exists for the SOC 2 framework. A research note from the Cloud Security Alliance confirms that no exposure draft, working group, or task force inside the AICPA has even proposed any of these yet.
The one official AICPA statement that touches AI comes from its Forensic and Valuation Services Executive Committee, and it's explicitly labeled as neither authoritative guidance nor a standard. By 2026, the AICPA had also put out guidelines and FAQs covering AI in federal tax practice. That guidance just tells practitioners how they might use AI tools in their own work. It says nothing about how to audit a client's AI system under SOC 2.
Into that vacuum, individual auditors have started building their own evidence requests, and the requests vary firm by firm. Some ask for training lineage: the dataset snapshot, the code commit, the hyperparameters, the approval that pushed a model into production. Some ask for per-inference logging: model version, a redacted prompt or a prompt hash, which tools got called, what the outcome was. Some ask about LLM subprocessor risk, so they push into a vendor's own compliance status and data-retention setup. Some ask for drift-monitoring dashboards paired with change-management tickets, so every model deployment can be shown to have been approved, tested, and reversible.
You won't find any of these categories anywhere in the AICPA's Trust Services Criteria. They're overlays that individual practitioners built on their own, with no central body comparing notes or setting a floor for what counts as real "AI-specific SOC 2" evidence.
Baker Tilly, which absorbed Moss Adams in a June 2025 merger, offers a useful look at what good-faith improvisation looks like at a sophisticated firm. Its published guidance recommends embedding AI-specific controls directly inside the existing five criteria, and it treats model development, validation, and monitoring as extensions of processing integrity and security. That's a thoughtful interpretation from a firm with real resources behind it. It's also just one firm's interpretation, not a standard, and nothing requires any other firm to follow it.
This improvisation is unfolding alongside a separate quality problem that makes it harder to trust any given report on its face. The AICPA's Assurance Services Executive Committee has already warned that "fast and easy" audit platforms are churning out templated, look-alike reports, a concern that predates AI. Then, in May 2026, the AICPA Peer Review Board issued a reviewer alert that told peer reviewers to apply added scrutiny to high-volume SOC 2 practices, backed by a structured monitoring and outreach program that launched June 1, 2026. Auditors are filling a real gap, and the broader system checking their work is simultaneously trying to clean house.
What a Clean Report Certifies
Picture two AI vendors, and each one holds a clean SOC 2 Type II report. They can have tested entirely different control sets from each other, because no common AI baseline exists for either auditor to test against. A buyer sitting across from both reports has no way to tell whether "processing integrity" meant model drift monitoring for one vendor and plain old data-processing accuracy for the other. The report doesn't say. It can't say, because the criteria behind it never defined the term precisely enough to force a consistent answer.
The assumption gap runs in both directions. Plenty of organizations believe certification automatically protects their AI systems, and they never check what scope their own auditor actually tested. Plenty of buyers assume a clean report from a vendor covers that vendor's AI capabilities specifically. What the criteria actually require is considerably less than either side imagines, so neither assumption holds up.
Auditor practice is tightening in one respect: static screenshots get challenged more often now, and auditors increasingly expect continuous-monitoring exports with real timestamps attached. That's a genuine improvement in rigor. It still doesn't touch the AI-specific gaps, because the underlying criteria still don't name shadow AI, vendor model risk, drift, or agentic accountability as things to test. Tighter scrutiny applied to the wrong checklist still produces the wrong answer.
The AI Trust OS paper names the commercial stakes clearly: the gap is operational, commercial, and regulatory at the same time. Enterprise procurement teams increasingly want real-time, empirical proof that a vendor's AI governance is mature. A SOC 2 report, however clean, does not provide that proof. For a risk reviewer staring at a vendor's report during due diligence, the whole thesis becomes personal at that moment: the report isn't lying to anyone, but it's silent on precisely the question that reader most needs answered.
The strongest objection to treating this as a framework failure
The best case against everything above goes like this: SOC 2 and ISO 27001 were built to be technology-agnostic on purpose, and that design choice has let them survive multiple technology generations without needing a rewrite each time. Bake in AI-specific criteria now, the argument goes, and the standard risks going stale the moment the technology shifts again, which in this field could be a year or two away.
It's a real argument, and it deserves a real answer rather than a dismissal. Technology-agnostic design works fine when auditors can take a new technology, apply professional judgment, and map its risks back onto criteria that already exist. That's how SOC 2 absorbed cloud computing. That's how it absorbed mobile. Agentic AI breaks that pattern in a different way: both frameworks' evidence-collection method assumes a human practitioner can manually gather screenshots and compile artifacts by hand. That assumption falls apart once the system spans five vendors, runs data through embedding models, and logs its own activity to observability platforms, and it does all that inside a single user request that takes seconds to complete. There's no screenshot that captures that.
The UNSW coverage matrix makes the point sharper still. None of the five major frameworks it tested, ISO 27001 included, address shadow AI, speed asymmetry, or the governance vacuum at the level of detail the problem demands. Technology-agnostic design is not just stretching a little to cover a new case here; it is structurally silent on the failure modes most likely to cause real harm.
Where structured responses are beginning to form
Three developments since late 2025 have moved the landscape forward, and none of them closes the gap on its own.
The Cloud Security Alliance's STAR for AI program is the clearest sign that transparency at enterprise scale is achievable. In November 2025, CSA recognized Microsoft and Zendesk as the first two organizations in the world to reach STAR for AI Level 2 certification. Earlier that October, Anthropic, Sierra, and Zendesk had already posted ISO 42001 certificates in the STAR Registry, showing that early adopters see commercial value in this kind of transparency. CSA's own agentic certification scheme targets initial auditor qualification requirements, pilot program enrollment, and registry integration, and it spans the second and third quarters of 2026. Even CSA acknowledges the program still needs an agentic-specific layer before it can address these risks on its own.
The AICPA's rulemaking calendar tells a slower story. The main attestation-standard activity on the books for 2026 covers proposed revisions to AT-C sections 105, 205, and 210, alongside new sustainability-specific sections 325 and 330. That work is aimed at sustainability assurance and general examination methodology. It doesn't mention AI. And it wouldn't take effect for engagements beginning before June 15, 2029, which leaves a gap measured in years, not quarters.
ISO 42001 fills part of the space differently. It requires structured inventories of AI systems, formally documented risk assessments, and demonstrable evidence that controls were actually implemented, and these requirements line up with AI-specific risk far more directly than SOC 2 or ISO 27001 ever did. It's a management system standard, not an attestation framework, so it sits alongside SOC 2 and ISO 27001 rather than replacing either one.
Even the broader risk-naming effort is fragmenting rather than converging. At least 74 AI risk taxonomies now exist across the field, and most of them stop at naming a risk without showing how to turn that name into a test run against a real system, with a measured outcome and a defensible grade. That's the same operationalization gap auditors run into inside SOC 2 and ISO 27001 itself: knowing the risk exists is not the same as having a way to check for it.
Meanwhile the window for consequence-free improvisation is narrowing. AICPA Peer Review scrutiny, in effect since June 2026, now applies to the ad hoc testing practitioners have been doing on their own. Auditors face more oversight for their improvised evidence requests at the exact moment the authoritative standard they're improvising around still hasn't arrived.
Continuous, telemetry-driven monitoring of the actual AI attack surface
Everything in the previous six sections points toward the same conclusion: better checklists will not fix this. The fix requires a different way of knowing what's true about an AI system in the first place, moving from organizations self-reporting what they believe is true to actually watching the system run.
The AI Trust OS framework frames this as a categorical shift, not an incremental upgrade to existing compliance tools. Four changes sit at the center of it: proactive discovery instead of reactive declaration, telemetry evidence instead of manual attestation, continuous posture instead of a once-a-year audit snapshot, and proof backed by the system's own architecture instead of trust placed in a policy document.
If you apply that shift to each blind spot named earlier, the shape of the answer becomes concrete. Shadow AI assets have to be found through observability signals, because a tool an employee picked up without telling IT will never appear on an asset register handed to an auditor, no matter how carefully that register gets compiled. You need continuous monitoring of vendor model risk across the full chain of LLM subprocessors, not just a single check at onboarding against a vendor's own SOC 2 report, because that report may never have tested the AI-specific controls that matter.
That's the honest shape of where things stand. SOC 2 and ISO 27001 still serve the purposes they were built for. They were never built to see an autonomous agent calling five vendors inside one request, and no amount of careful auditing inside the existing criteria changes that fact, because the frameworks are just answering a question that predates the one enterprises are actually asking now.
Sources
- The SOC 2 AI Gap: Auditors Improvise Where AICPA Has Not Acted
- AI Trust OS -- A Continuous Governance Framework for Autonomous AI Observability and Zero-Trust Compliance in Enterprise Environments
- AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis
- The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
- Evolving SOC 2 reports for AI controls


