← All news
Press · August 31, 2026 · 11 min read

AI Agent Decision Auditability: The Log Shows What the Agent Did, Not What It Chose Between

AI Agent Decision Auditability: The Log Shows What the Agent Did, Not What It Chose Between

Glean gave agents their own identity and audit log. The log records which source the agent cited, never which valid source it discarded.

On 26 August 2026, Glean announced its “independent agents”: autonomous agents with their own identity, permissions provisioned separately from any single user’s, and their own audit log. Seven months earlier, on 22 January 2026, Singapore’s IMDA published a governance framework for agentic AI that asks for precisely this — that every agent carry a verifiable digital identity and a trace of who authorised it to act.

Both moves point the same way, and that is good news: the autonomous agent stops borrowing an employee’s account and becomes an identifiable entity whose actions are recorded. For a CDO in a regulated environment, that is the entry condition for any serious deployment.

It is not sufficient. That kind of log records the agent’s action, the identity it acted under, the timestamp, the resource it read. It does not record what the agent chose between. When two equally valid documents in the corpus contradict each other, no log line says an arbitration took place, because the agent never experienced one: it ranked, it cited the top-scored source, it acted. Call this the arbitration gap: the distance between the trace of what an agent did and the trace of what it set aside.

The test takes half a day and needs no tooling. Take the three most-consulted procedures in your document estate. Ask each business owner to name the version that is authoritative today, and the date the previous ones were withdrawn from circulation. How often that answer comes back without hesitation is your baseline.

This blog has already covered the governance of the agents themselves (5 August, on autonomy frameworks) and the data governance obligations of Article 10 of the AI Act (6 July). The subject here sits downstream of the first and upstream of the second: reconstructing, six months later, a document decision an agent made on its own. Organisations typically answer this in one of two ways, and the two look more alike than they appear. Either a home-grown process, where business owners review their documentation before the corpus is opened to AI. Or the tool already on the balance sheet: an AI governance platform, or an identity layer for agents, logging accesses and actions in fine detail. The first does not scale to a real document estate. The second governs the agent, never the document estate the agent consumes.

What an agent log records, and what it structurally cannot

An agent audit log is a register of technical events. It answers: which agent, under which identity, with which permissions, at what time, on which resource, with what outcome. Well maintained, it lets you replay a chain of actions and attribute responsibility to a human owner.

Take it in its best form, the one serious vendors ship today: a log that captures not just the action but the identifier and version number of the document cited, timestamped and retained for a long period. That is a good log, and it genuinely changes post-incident investigation.

The boundary is right there, and it is sharp: its reach stops at the source that was cited. The existence of a valid source that was set aside appears nowhere in it. It shows that the agent cited procedure PROC-4412-v3. It does not show that an earlier version, never withdrawn from circulation, carried a different threshold; that the March internal memo, published in another workspace, said the opposite; or that both documents were technically valid, unexpired and correctly indexed. The log is complete from the standpoint of execution, and silent from the standpoint of document authority.

This is a layering problem. Execution traceability is built at the moment the agent acts. Document authority traceability is built beforehand, on the estate itself, once and then continuously. No amount of downstream logging can reconstruct information that was never produced upstream.

And the asymmetry matters. A human faced with two diverging answers hesitates, asks, escalates: doubt is a detection mechanism, imperfect but real. An agent running continuously and initiating its own actions has no such mechanism. It arbitrates on a relevance score, without perceiving that it is arbitrating at all.

The multiplier: from one wrong answer to N unreplayable decisions

As long as a human asks the question, a document contradiction costs one wrong answer, at one moment, to one person who can doubt and verify. The damage is discrete and locatable.

An autonomous agent changes the unit of account. The same undetected conflict produces one decision per occurrence, every day, across every affected case, with perfect consistency. Each decision is individually plausible, logged, attributed to an identified agent. None is replayable, because nothing in the trace indicates that a valid alternative existed. This is a problem of scale and silence more than one of quality.

Information retrieval research has a name for the case. Work on knowledge conflicts distinguishes conflict between the model’s parametric memory and the retrieved context from inter-context conflict: between several documents retrieved at once. The ConflictQA benchmark, presented at SIGIR 2026, notes that this second case has remained largely under-studied even though it is the most common in a real enterprise corpus, and measures that explicitly telling the model which type of conflict is present markedly improves the faithfulness of its reasoning, by roughly twenty points under its protocol. The magnitude belongs to that protocol, not to your estate; the mechanics do transfer. Conflict detection is a separate signal, and it has to be produced on the estate before the model is ever called. If nobody produced it, the model will not invent it, and the log will not record it.

What traceability has to produce to be defensible

On the regulatory side, the European calendar has loosened without the substance of the requirements changing. Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026 and in force since 27 July, defers compliance for standalone high-risk systems under Annex III to 2 December 2027, and for AI embedded in products already covered by sectoral legislation (Annex I) to 2 August 2028.

That deferral releases nothing on the perimeter already in force. Transparency obligations have applied since 2 August 2026. And on the high-risk perimeter, the articles say what they were voted to say, with a division of roles worth keeping in mind, because it determines what is left for you once your vendor has done its part. Article 12 is a design requirement on the provider of the system: it must enable automatic recording of events across the lifecycle. Articles 19 and 26 are deployer obligations, meaning yours: retain those logs for at least six months, and assign the human oversight required by Article 14 to people with the necessary competence and authority. A vendor can hand you a log that is impeccable under Article 12 without a single one of your document decisions becoming reconstructable. The deferral buys preparation time, and nothing else. For an organisation whose systems move into the high-risk perimeter in December 2027, the deadline that matters is getting the document estate in order, and that is counted in quarters.

Effective human oversight of a document decision assumes the person supervising can answer one plain question: which version of which document was authoritative on the day the agent decided, and where is the trace of the arbitration? An execution log answers the first half. The second half exists only if the contradictions in the corpus were identified, arbitrated and dated beforehand, by named business experts.

That is the remit of a Document Knowledge Platform (DKP): the document quality layer upstream of AI systems, in three movements. Govern, establishing ownership, authority and lifecycle for every document. Clean, detecting and resolving anomalies, diverging duplicates, obsolescence and contradictions. Activate, opening the corpus to AI systems only once the first two hold. This is neither one more enterprise search engine nor a model governance platform: it is the layer that runs before them, on the estate they consume. That position is stable and does not get redefined by the vendor announcement cycle or by the analysts’ publication calendar. The document layer precedes the engine, whichever engine the quarter brings.

One point of reassurance on method, handled here rather than in an appendix because this is where the mechanics are described: the diagnostic covers document content only, never conversation transcripts, usage logs or user telemetry. The ingestion perimeter is contractual, and documents are not reused to train models.

Count the contradictions before granting the right to act

The measurement is available, and it is boring in the best sense: it consists of counting.

In an anonymised case in the energy and industrial sector, an initial diagnostic identified 398 conflicts in the document estate; resolving them raised answer reliability by more than 90%. Those 398 conflicts existed before any AI deployment. None was visible without counting, because every document, taken on its own, was valid.

At a CAC 40 industrial group, the same exercise showed that 32% of the base consisted of diverging duplicates: surfaced in two weeks, resolved in six. A diverging duplicate is the case most hostile to an agent’s auditability, because both versions pass every formal control. The relevance score picks one, and nobody is told.

The resulting order of operations is simple. Audit: count the contradictions, diverging duplicates and expired documents in the corpus the agent will be allowed to act on, before you grant that right. Clean: have every conflict arbitrated by the competent business expert, and keep the dated trace of that arbitration, which is the only artefact that will make the agent’s decision reconstructable six months later. Monitor: repeat the measurement continuously, because an estate degrades with every publication, and because autonomous agents now author documents themselves.

Before you give an agent an identity, permissions and the right to act, count the contradictions in the corpus it will act on. It is the one question its audit log will never be able to answer for you.

That is precisely the deliverable of the entry engagement: the count of contradictions and diverging duplicates in your estate, documented one by one and routed to the competent business expert, with the twenty most critical cases handled first. Joint sign-off by the business Document Owner and the CISO or DPO, never IT alone.

Frequently Asked Questions

Isn’t an AI agent audit log enough to satisfy an auditor?

It answers the question of action: which agent, under which identity, on which resource, at what time. It does not answer the question of authority: was the cited document the authoritative one, and did a contradicting version coexist. That second piece of information has to be produced on the document estate upstream; it cannot be derived from the log.

We already have an AI governance platform and identity management for agents. What is missing?

Those layers govern the agent: its identity, permissions, actions and lifecycle. They do not govern the document estate it consumes. The classic blind spot is the diverging duplicate: two valid, unexpired, correctly indexed documents that say the opposite. No access control detects it.

Does the AI Act deferral to December 2027 let us wait?

Regulation (EU) 2026/1744 defers the Annex III high-risk deadline to 2 December 2027 and the Annex I deadline to 2 August 2028. Transparency obligations, for their part, have applied since 2 August 2026. On the high-risk perimeter the substance is unchanged: event recording, log retention, effective human oversight. Getting a document estate in order is counted in quarters, not weeks.

Who should sign off on a corpus diagnostic inside the organisation?

The business Document Owner, who owns both the pain and the arbitration decision, jointly with the CISO or DPO as perimeter guardian. Not IT alone: IT integrates, it is not the document authority. Business experts remain the decision-makers on each conflict; the diagnostic prepares and routes, it does not decide for them.

What exactly is in scope from a confidentiality standpoint?

Document content only: no conversation transcripts, no usage logs, no user telemetry. The ingestion perimeter is contractually defined, and documents are not reused for model training.

Sources


Where to Go From Here

K-AI Corpus Diagnostic — 10 business days on your document estate, full report of the 20 most critical anomalies, money-back guarantee if no meaningful anomaly is found. To count the contradictions in a corpus before granting an agent the right to act on it, reach the K-AI team: contact@k-ai.ai. The scope of every diagnostic is validated jointly by the business Document Owner and the CISO/DPO, never by IT alone.

K-AI already works with CMA CGM, Veolia, PwC, BNP Paribas, TotalEnergies and CEVA Logistics. Partners: AWS, Snowflake, Microsoft, Wavestone, Devoteam.

And in your organization, what does your document estate look like?

30 minutes with a founder. We audit a sample of your documents for free and show you exactly what K-AI detects.

Book a demo → Read other articles