← All news
Press · September 2, 2026 · 13 min read

Cost per task in enterprise AI: the metric that turns document quality into a budget line

Cost per task in enterprise AI: the metric that turns document quality into a budget line

Gartner: only 22% of enterprises have scaled AI. Around 60% of an agentic task's cost goes to rework — a corpus problem, not a model problem.

On 1 September 2026, Gartner published an uncomfortable number: only 22% of organisations have successfully scaled AI across multiple business units, while 85% of functional leaders plan to increase their AI spending this year. Three weeks earlier, the same firm noted that in 2026, for the first time, inference spending overtakes training spending.

Read together, these two facts say the same thing. AI spending is leaving the project column and entering operations. And once spending becomes operational, it gets measured per unit of use, which means per task. Cost per successful task is therefore settling in as the metric executive committees will use to arbitrate, and it will be on the table of every investment committee over the next eighteen months.

The argument of this article is straightforward: that metric is not a model metric, nor an infrastructure metric. It is a document corpus metric. The dominant cost line in an agentic task is the rework triggered when the first answer does not hold. And in most of the cases we see in the field, what makes a first answer fail is a property of the document estate, not of the engine reading it.

This is written for a specific reader: someone who already runs an assistant, a chatbot or an agent in production, or is scaling one right now, and whose inference bill has overshot the estimate in the investment case. If your deployment is still a closed pilot, the argument will hold in six months. It does not bite yet.

The order of magnitude worth carrying into the rest of this piece comes from a measured case, detailed further down. On a customer documentation estate of roughly 500 web pages, 19% of pages carried an ambiguity or a contradiction, and 53% of cases were resolved in three weeks, prioritising the most critical, by one to two business experts working half a day to one day per week. It is the ratio between those two terms that interests a CFO.

A word on how this differs from our 21 August piece, which dealt with governing a corpus that grows under the pressure of AI-generated content. The subject here is different: it concerns what that corpus costs on every single execution, rather than how large it has become.

What cost per task actually measures

One order of magnitude before the mechanics. Work published in July 2026 by McKinsey QuantumBlack on the economics of agentic workflows puts around 60% of an agentic task’s cost in answer refinement: checks, retries, re-prompts, validation loops that fire when the first pass fails to clear the expected quality bar. Estimates of the token overhead of an agentic workflow compared with a single conversational exchange vary too much across protocols and scopes for any single order of magnitude to be defensible, and we refrain from quoting one. They do converge on one point: the centre of gravity of cost sits in the rework. Gartner, for its part, predicted as early as 25 June 2025 that over 40% of agentic AI projects would be cancelled before the end of 2027, citing cost overruns as a leading cause.

A rework loop fires for very concrete reasons. The agent queries the corpus, gets two equally plausible answers from two equally valid documents, fails to decide, reformulates, widens its search, calls a tool again, produces a hedged answer, or picks one according to a relevance score, which is cheaper now and far more expensive later. On every execution. On every affected question. This is what we call re-verification debt: the recurring overhead an AI system pays because its corpus does not let it conclude on the first attempt. It is counted the only way our trade knows how to count, by enumerating the contradictions in the corpus. Same exercise we run on our clients’ document estates, with a budget consequence attached.

Two reflexes will show up in committee, and they are worth naming immediately. The first is to treat this as a control problem: cap tokens, put consumption monitoring in place, charge the spend back to the business units. The second is to assume the tool already purchased covers the ground, whether that is an AI observability platform, a FinOps layer, or the data catalogue already in service. Those tools do exactly what they were built for: they tell you how much you spend, where, and by whom. None of them tells you why the task cost more than forecast, because the cause does not live in the telemetry. It lives in the documents.

There is a test that requires no tool and no project, and takes half a day. Take the twenty questions most frequently put to your internal assistant. For each one, count how many documents in the estate claim to answer it, and how many of those carry authority. Every question with more than one authoritative answer is a re-verification loop you pay on every execution, every day, with no budget line bearing its name.

Why document quality changes register

Until now, document quality was defended in committee with two arguments: answer reliability and compliance. Both solid, and both outside the operating budget. They belong to risk, therefore to provisioning, therefore to slow arbitration.

Cost per task moves the argument into a different column. A contradiction in the corpus becomes recurring inference consumption, on top of the wrong-answer risk it already carried. It has a frequency, every execution, and a volume, the number of calls affected. In other words it becomes a variable cost line, and document remediation becomes an investment with a computable return rather than an insurance policy.

That still needs to be attributable, and a CFO will put the question plainly: what are you going to show me that goes down? The calculation uses three terms available in any deployment. One, the call frequency of each intent, which your usage logs already hold. Two, the number of competing authoritative sources on that intent, which the corpus audit produces. Three, the observed cost gap between an execution with rework and one without, which your inference platform bills line by line. The product of the first two gives the volume of rework attributable to the corpus; the third converts it into currency. None of these three requires a new tool. The one almost always missing is the second.

This matters for a CDO defending a budget. Gartner notes in its 1 September survey that roughly 11% of organisations are entirely unaware of what their function spent on AI in 2025, and that high performers are precisely those tracking the return of each initiative by outcome category. The corpus is one of the few line items where the relationship between a correction and a unit saving is directly observable, because the correction is dated and the call volume is known.

One boundary, to avoid a misreading. A document quality and governance layer sells on the reliability and defensibility of answers, not on cost optimisation, and document quality does not thereby become a FinOps topic. Cost per task is the channel through which an old problem becomes visible this year, in a budget column the committee reviews every month. The inference saving is its quantifiable consequence.

This position does not depend on any analyst artefact. It is not redefined by the rhythm of Magic Quadrants, Hype Cycles or benchmark leaderboards: the document layer precedes the engine, whatever the engine and whatever the quarter.

What a field measurement shows

TotalEnergies Retail Power & Gas runs a customer-facing chatbot in production, serving consumers, professionals, businesses and local authorities, fed by roughly 500 pages of official web documentation. In the diagnostic run with K-AI, 19% of those pages needed correction, including contradictions impossible to spot by eye at that scale. 53% of cases were resolved in three weeks, prioritising the most critical, with deliberately limited resources: one to two business experts, half a day to one day per week. The chatbot’s accuracy rate was measured before and after remediation. These figures apply to that specific scope at that specific moment and do not generalise.

What this case shows about cost is the asymmetry stated in the introduction: one to two business experts, half a day to one day per week over three weeks, against a system that was paying for that ambiguity on every customer conversation. That gap is where a CDO’s room for manoeuvre sits.

A methodological point that matters as much as the result: business experts remain the decision-makers. K-AI produces the diagnostic, ranks the cases and indicates which expert each anomaly should be routed to; the document owners decide. The technical process ingests document content only, no conversation transcripts, no usage logs, no telemetry, within a contractually defined ingestion scope, with no reuse of content for model training.

Three decisions if your systems are already in production

Three concrete trade-offs face the reader described in the introduction, the one whose inference bill has overshot the estimate.

Your next investment case documents the state of the corpus, or it cannot defend its estimate. An AI business case today documents the use case, the architecture, the vendor and compliance. It almost never documents the state of the document estate the system will run against. Yet that is the variable that decides whether your forecast cost per task will resemble the observed one. Add a line to it: number of intents with competing authoritative sources across the exposed scope.

Your vendor comparison has to happen at constant corpus, or it measures nothing useful. The financial stake is now explicit: on 1 July 2026 Gartner put $234 billion of enterprise application software spend at risk from the agentic shift. On an estate that contradicts itself, the unit cost gap between two platforms mostly describes how each one reacts to noise, not the value it will bring you. Vendor-published cost-per-task benchmarks are, as of today, not externally reproducible: they do not replace a measurement made on your own corpus.

Your remediation has to be recurring, because the spend it reduces is. A one-off cleanup lowers cost per task for a quarter. The corpus keeps living: new versions, new procedures, generated content. The inference bill arrives every month. Setting a single correction against a recurring cost is a bad trade, and it is the most common mistake we see after a successful first diagnostic.

This is the logic of a Document Knowledge Platform (DKP): govern the document estate (Govern, covering ownership, authority, lifecycle), detect and treat anomalies, duplicates, obsolescence and contradictions (Clean), then activate the corpus for AI systems only once those two steps hold (Activate). A DKP replaces neither a data catalogue, nor a data and AI governance platform, nor an ECM, nor an enterprise search engine, nor an agentic knowledge layer, nor a FinOps layer. It runs upstream, on the estate all of those consume, and it runs continuously.

On the regulatory side, one note to prevent a frequent misreading. As the firms tracking the text set out, White & Case among them, Regulation (EU) 2026/1744, the Digital Omnibus, in force since 27 July 2026, has deferred several obligations: Annex III and Article 6(2) to 2 December 2027, Annex I and Article 6(1) to 2 August 2028. That deferral suspends nothing already applicable within the scope in force: transparency obligations, governance of training and context data, and human oversight of systems already in production remain enforceable. A calendar that slips is not permission to deprioritise.

Conclusion: audit, clean, monitor

Cost per successful task is becoming the question your investment committee will ask before approving the next deployment. Answering it means knowing how many times, across your estate, two valid documents claim to answer the same question. A better model does not produce that information.

The sequence is the same one we apply on every engagement. Audit: count the contradictions, divergent duplicates and obsolescence across the estate actually exposed to AI systems. Clean: have document owners settle the cases, starting with the most frequently asked questions. Monitor: measure continuously, because the corpus keeps moving and the inference spend never stops.

Frequently Asked Questions

Isn’t cost per task primarily a question of model or architecture choice?

The model choice sets the unit price of a call. The number of calls needed to reach a result is decided elsewhere: in the system’s ability to conclude on the first attempt, therefore in the clarity of the corpus it queries. A cheaper model on a self-contradicting corpus can end up more expensive per successful task than a pricier model on a clean one.

Our CFO will ask which line goes down after the diagnostic. What do we show?

The volume of rework attributable to the corpus, computed from three terms you already own: the call frequency of each intent (your usage logs), the number of competing authoritative sources on that intent (the corpus audit), and the cost gap between an execution with rework and one without (your inference bill). The diagnostic produces the second term, which is generally the only one missing. The reduction then reads off the inference bill for the scope concerned, at comparable usage volume.

How do we estimate our re-verification debt without launching a project?

Take the twenty most frequent questions put to your internal assistant and count, for each, how many documents in the estate claim to answer it authoritatively. Every question with multiple answers is a loop paid on every execution. The test takes half a day and needs no tooling.

Our AI observability platform already tracks our costs. What does that not give us?

It gives you the measurement: how much, where, by whom. The documentary cause stays out of reach of telemetry, which observes system behaviour rather than document content. The two are complementary: one flags the cost drift, the other identifies its origin in the estate.

What document scope is ingested during a diagnostic, and under what guarantees?

Document content only, no conversation transcripts, no usage logs, no telemetry, within a contractually defined ingestion scope, with no reuse of content for model training. Scoping is validated jointly with the business Document Owner and the CISO or DPO, never with IT alone.

Does the AI Act deferral buy us time on this?

The Digital Omnibus deferred Annex III to 2 December 2027 and Annex I to 2 August 2028. It does not suspend obligations already applicable: transparency, governance of context data, and human oversight of systems in production. Economically, the deferral changes nothing: the inference spend is running today.

Sources


Where to Go From Here

K-AI Corpus Diagnostic — 10 business days on your document estate, full report of the 20 most critical anomalies: contradictions between valid documents, divergent duplicates and obsolescence across the scope actually exposed to your AI systems. Money-back guarantee if no meaningful anomaly is found. To put a number on your re-verification debt before the next investment committee, reach the K-AI team: contact@k-ai.ai. The scope of every diagnostic is validated jointly by the business Document Owner and the CISO/DPO, never by IT alone.

K-AI already works with CMA CGM, Veolia, PwC, BNP Paribas, TotalEnergies and CEVA Logistics. Partners: AWS, Snowflake, Microsoft, Wavestone, Devoteam.

And in your organization, what does your document estate look like?

30 minutes with a founder. We audit a sample of your documents for free and show you exactly what K-AI detects.

Book a demo → Read other articles