Agentic AI · Evidence Intelligence & Control Validation

Sufficient Evidence Is Still Audit's Biggest Gap.

The PCAOB's 2024 inspection results, published March 2025, found that 39% of inspected engagements still had a Part I.A deficiency — the auditor failed to obtain sufficient appropriate evidence. Evidence intelligence closes that gap earlier: an agent collects, matches, and scores control evidence continuously, before an auditor or regulator has to ask twice.

Crest.Digital Editorial August 30, 2026 11 min read Agentic Risk & Continuous Assurance

Ask any SOX programme lead or internal audit director what actually eats the weeks before a testing cycle closes, and the answer is rarely "figuring out what the controls should be." It's evidence: chasing screenshots from a control owner on leave, confirming a sign-off actually came from someone with the authority to give it, checking whether last quarter's file got quietly reused, and reconciling what a document claims against what the ERP or general ledger actually shows. None of that work tests whether the control design is sound. All of it determines whether the organization can prove the control ran.

That distinction — between a control existing and a control being provably evidenced — is precisely where audits keep failing, and the data confirms it isn't improving as fast as budgets are growing. The PCAOB's most recent inspection results, published in March 2025 and covering 2024 audits, put a hard number on a problem most control owners already feel every quarter: evidence sufficiency remains the single most common reason an audit opinion gets flagged.

Evidence intelligence is the practice this article covers: an AI agent that collects control evidence directly from the systems where it already lives, matches each item to the specific control requirement it is meant to satisfy, checks it for completeness, timing, ownership, and duplication, cross-references it against underlying system data, and assigns it a confidence score before a human reviewer, internal auditor, or external auditor ever asks for it. Nothing here decides what "sufficient" means for a given control — that judgment stays with the control owner, internal audit, and the auditor. What changes is how much of the checking that precedes that judgment gets done automatically, continuously, and consistently.

How much of your evidence library would survive a PCAOB-style sufficiency test today?

See how Crest.Digital's Agentic Risk & Continuous Assurance practice extends AI across evidence collection, control validation, and audit-readiness workflows — not just periodic manual sampling.

Explore Crest Intelligence

The Evidence Sufficiency Gap Is Getting Wider, Not Narrower

The PCAOB's Spotlight on 2024 inspection activities found the aggregate Part I.A deficiency rate — audits where the firm failed to obtain sufficient appropriate evidence to support its opinion on the financial statements or on internal control over financial reporting — was 39% across all inspected engagements, an improvement from 46% in 2023 but still meaning roughly two in five inspected audits had an evidence shortfall serious enough to be formally cited. Big Four firms improved to 20% from 26%. Non-affiliated firms inspected annually still sat at 52%. Insufficient testing of estimates, and of the underlying data and reports used to support audit conclusions, remained among the most recurring root causes across firm categories.

📋
39% of Audits, 40 Systems, 17% Automated The PCAOB's 2024 inspection results (published March 2025) found a 39% aggregate Part I.A deficiency rate for insufficient audit evidence. In the same reporting window, KPMG's 2025 SOX Survey found average SOX programme budgets rose 44% from FY2022 to FY2024 and average in-scope systems more than doubled from 17 to 40 — yet the share of fully automated controls actually declined, from 21% to 17%. Scope and cost are both climbing while automation share moves the wrong way.

Read together, those two data sets describe the same underlying mechanism from opposite sides. The PCAOB numbers show evidence sufficiency failing often enough to be a persistent, named category of audit deficiency. The KPMG numbers show why: the average SOX programme now spans more than double the systems it did two years ago, budgets have grown 44% to keep pace, and the manual-testing burden that scope growth creates is landing on teams whose share of automated controls is shrinking rather than growing. More systems, more evidence requests, less of it automated — that combination is exactly the setup in which a screenshot gets pulled from the wrong period, an approval gets rubber-stamped without verification, or last quarter's file gets reused because nobody had time to notice it was identical.

This is not a story about control owners cutting corners. It's a story about evidence-sufficiency checking — confirming period, ownership, approval, completeness, and non-duplication for every piece of evidence across a growing system footprint — being fundamentally a volume problem that manual review was never built to scale with.

What Evidence Intelligence Actually Does

In practice, the agent connects to the places evidence already lives — email, SharePoint or document management, the GRC or SOX platform, and the underlying ERP, procurement, payroll, or CRM systems the controls actually run in — and pulls the evidence associated with each control on its testing schedule. It then matches that evidence against the specific control requirement: the right control ID, the right testing period, the named accountable owner, and the required approval chain.

From there, the checking gets specific. Does the evidence actually cover the period being tested, or was a prior period's file attached by mistake? Does the approval on the document come from someone who currently holds the authority to give it, or from a role that changed hands two quarters ago? Is the evidence complete against what the control procedure specifies, or is a required attachment missing? Has this exact file, or a suspiciously similar one, already been submitted for a prior testing cycle? And where the underlying system can be queried directly, does what the evidence claims — a reconciliation total, an approval count, a threshold — actually match what the system of record shows? Each check produces a discrete finding, and the aggregate becomes an AI-generated evidence-confidence score that tells a reviewer, at a glance, which controls are genuinely audit-ready and which need attention before the auditor asks.

This is deliberately the layer below agentic internal audit, which uses evidence quality as one input into risk-based sampling and audit programme design, and it sits downstream of policy-to-control intelligence, which defines in the first place what evidence a given control should require. Evidence intelligence is the layer that actually tests whether what got submitted meets that bar, continuously, rather than during a periodic sample taken once or twice a year.

Still verifying evidence sufficiency manually across a growing system footprint?

Crest.Digital designs customized AI agents that collect control evidence from your existing systems, check it for completeness and authenticity, and score it before your next audit cycle — connected to the GRC, ERP, and document platforms you already run.

What the Standards Already Expect

Evidence sufficiency is not a new bar being set by AI vendors — it is the specific standard the PCAOB has always applied, and its newly adopted AS 1215 on audit documentation, effective for audits of fiscal years ending on or after December 15, 2026, sharpens the requirement further by tightening what counts as adequate documentation of the evidence obtained and the conclusions reached. The COSO Internal Control–Integrated Framework makes the same point structurally in its Information and Communication component: relevant, quality information has to be available on a timely basis, which is a documentation and evidence problem as much as a data problem.

ISACA's COBIT guidance on continuous assurance treats evidence collection as an ongoing activity rather than a periodic exercise for exactly this reason — a control tested once a quarter only proves itself for the moment it was sampled, while continuously validated evidence proves it across the whole period. The IIA's Global Internal Audit Standards, effective from January 2025, likewise place explicit weight on the quality and sufficiency of evidence internal audit relies on to support its own conclusions, not only the conclusions of the functions it reviews.

KPMG's 2025 SOX Survey adds the practitioner-side confirmation: as in-scope systems and testing volume expand faster than automation, the manual evidence-review bottleneck compounds rather than resolves itself on its own. Read alongside the PCAOB's persistent 39% deficiency rate, the direction from regulators, standard-setters, and practitioner surveys is consistent — sufficiency of evidence is the standard that keeps getting missed, and it is a scale problem before it is a judgment problem.

An 8-Point Framework for Evidence Intelligence

Standing up evidence intelligence doesn't require replacing how a control owner or audit team already thinks about evidence. It formalizes and continuously runs the same checks a thorough manual reviewer already performs — just applied to every control, every cycle, without the volume constraint.

1

Evidence Collection & Ingestion

Pull evidence directly from email, SharePoint, document repositories, and the ERP or business systems the control runs in, rather than waiting on manual upload.

2

Control-Requirement Matching

Match each item of evidence to the exact control ID and testing requirement it is meant to satisfy, flagging evidence that can't be matched at all.

3

Period, Owner & Approval Verification

Confirm the evidence covers the correct testing period and carries approval from someone who currently holds the authority to give it.

4

Completeness Checking

Check the evidence package against every attachment and data point the control procedure specifies, flagging anything missing.

5

Recycled & Duplicate Evidence Detection

Detect evidence that matches or closely resembles a prior cycle's submission instead of reflecting the current testing period.

6

Cross-Validation Against System Data

Where possible, check what the evidence claims against the underlying system of record — reconciliation totals, approval logs, thresholds.

7

AI Evidence-Confidence Scoring

Aggregate every check into a single confidence score per control, so reviewers can prioritize the weakest evidence first.

8

Escalation & Audit-Trail Maintenance

Escalate insufficient evidence to the control owner automatically, and log every check performed for a complete, defensible audit trail.

Points five and seven tend to draw the most scrutiny from audit committees and external auditors alike, so it's worth being direct about them: the agent flags a duplicate or a low-confidence score — it does not decide that a control has therefore failed. That determination, and any resulting remediation, stays with the control owner, internal audit, or the auditor, working from a complete and consistent set of findings instead of whatever a manual spot-check happened to notice.

Building the Programme: A Six-Step Delivery Playbook

Crest.Digital positions this work as a configurable "Risk Automation Pod" rather than a bespoke software build — a shared underlying stack of integration connectors, scoring engine, evidence repository, and dashboards, customized around a specific organization's control population, system landscape, and existing GRC or SOX tooling.

The Discover → Design → Connect → Deploy → Validate → Transfer Model

  • Discover: Inventory in-scope controls, the systems and repositories where evidence for each lives, and the control types where evidence gaps have recurred.
  • Design: Define evidence-sufficiency rules per control type, confidence-scoring logic, duplication thresholds, and human escalation checkpoints.
  • Connect: Integrate with email, document management, the GRC or SOX platform, and the ERP or business systems generating the underlying data.
  • Deploy: Run the agent across a selected control population, generating confidence scores and routing gaps to the named owner — starting with the highest-risk controls.
  • Validate: Compare the agent's assessments against a manual review from internal audit or the external auditor, and set accuracy thresholds before expanding.
  • Transfer or manage: Hand the workflow to the internal audit or SOX team, or continue as a Crest.Digital-managed service.

Starting with the control population that carries the highest inherent risk — revenue recognition, manual journal entries, or access-provisioning controls are common first choices — lets the validate step do genuinely useful work: comparing the agent's confidence scores against what an experienced reviewer would have concluded manually, before the organization leans on the scores for anything audit-facing.

Where the Agent Stops and Judgment Stays Human

The reasonable question every SOX lead or audit director eventually asks is how much of this can be trusted without a person checking every file. Evidence intelligence is built around a specific division of labor rather than removing judgment from the process. Agents are genuinely strong at three things here: checking evidence against defined rules — period, owner, approval, completeness, duplication — with total consistency across every control and every cycle; cross-referencing evidence against underlying system data at a scale no manual reviewer could match; and surfacing every low-confidence item continuously, instead of only at the point a periodic sample happens to catch it.

What the agent doesn't do is decide whether a flagged gap is material, whether a compensating control covers a genuine shortfall, or how a control should be redesigned if its evidence keeps falling short. Those calls require organizational context, risk appetite, and accountability that stay with the control owner, internal audit, and the external auditor — which is why every flagged exception routes back to a named human before it becomes a remediation decision, not after.

This human-in-the-loop discipline runs through Crest.Digital's wider agentic risk practice. If an agent flags a control as insufficiently evidenced, the reviewer needs to be able to reconstruct exactly what was checked and why — the same audit-trail requirement covered directly elsewhere, and the accountability question that follows — who signs off when an agent's scoring feeds a compliance or audit decision — is addressed in the companion piece on GRC accountability. Evidence-confidence scores also become a direct input into risk-based sampling and continuous assurance, letting audit teams prioritize the controls with the weakest evidence rather than sampling on a fixed schedule regardless of where the actual risk sits.

The underlying thesis is consistent with what Crest.Digital has argued across its TPRM work — a document proves what was submitted, not what's actually true. In a SOX or audit programme, that gap shows up as a screenshot from the wrong period, an approval that no longer reflects who holds authority, or a file recycled from last quarter. The fix is the same: keep the accountable reviewer's judgment intact, and give it a complete, current, scored picture of the evidence to work from — instead of whatever a periodic manual pass happened to catch before the auditor asked.

Frequently Asked Questions

Evidence intelligence is the use of an AI agent to collect control evidence directly from the systems where it already lives — email, SharePoint, document repositories, and the ERP or business application the control actually runs in — and match each piece of evidence to the specific control requirement it is meant to satisfy. The agent checks that the evidence covers the correct testing period, carries the right owner and approval, and is complete against what the control procedure calls for; flags evidence that looks recycled or duplicated from a prior cycle; cross-checks it against underlying system data where possible; assigns it an AI-generated confidence score; and escalates anything that falls short, so a human reviewer sees the exception before an auditor does.

A GRC repository is mostly a filing cabinet — it stores whatever a control owner uploads and shows a checklist of what's outstanding. It generally does not verify that a screenshot actually covers the right month, that an approval signature belongs to someone with the authority to give it, that the same evidence wasn't reused from last quarter, or that what the document claims lines up with what the underlying system actually recorded. Evidence intelligence sits on top of that repository and does the verification work a control owner or auditor would otherwise have to do manually, at a consistency and speed manual review can't match across hundreds of controls.

The PCAOB's Spotlight on 2024 inspection activities, published in March 2025, found that the aggregate Part I.A deficiency rate — engagements where the audit firm failed to obtain sufficient appropriate evidence to support its opinion — was 39% across all inspected firms, down from 46% in 2023 but still meaning roughly two in five inspected audits had an evidence shortfall serious enough to be cited. Big Four firms improved to a 20% deficiency rate from 26%, while non-affiliated firms inspected annually still sat at 52%. Insufficient testing of estimates, underlying data, and reports used to support audit conclusions remained among the most recurring root causes.

No. The agent's role is to check evidence against defined completeness, timing, ownership, and duplication rules, cross-reference it against system data where available, and produce a confidence score and a specific list of what's missing or inconsistent — then route that to the control owner or the audit team. Whether a given gap is material, whether a compensating control covers it, and how to remediate a genuinely deficient control remain judgment calls for the control owner, internal audit, or the external auditor. The agent's output is a consistent, evidenced starting point for that judgment, not a replacement for it.

Continuous controls monitoring tests whether a control's underlying transactions behave the way they should — flagging exceptions like duplicate payments or approval bypasses as they happen. Evidence intelligence operates one layer above that: it verifies whether the documentation proving a control ran is itself sufficient, current, and genuine. Agentic internal audit then draws on both — using the evidence-confidence scores and monitoring exceptions to build risk-based sampling, draft observations, and prioritize which controls need audit attention first. The three work together: one tests behavior, one tests proof, and one turns both into an audit programme a human auditor still directs.

Evidence Intelligence Control Validation SOX Compliance Automation Continuous Controls Monitoring Agentic Risk & Continuous Assurance